The open agent stack: model, router, and working memory
Three releases from NVIDIA and IBM Research move the search for efficient agents from an open model to call routing and selectively retrieved guidance from past runs.
Artificial Intelligence··Evening
A training package alongside the model weights
NVIDIA released Nemotron 3.5 Lightning, a mixture-of-experts model with 30 billion parameters, under the OpenMDW-1.1 license together with its weights, training data, and training recipes. The model activates 3 billion parameters for each token. NVFP4 and BF16 checkpoints are available through Hugging Face, ModelScope, and NVIDIA's own access points. The company says the model can be served through vLLM, SGLang, and TensorRT-LLM and can also run with Ollama, llama.cpp, LM Studio, and Unsloth. In NVIDIA's own measurements, it reaches up to 4 times the output speed of similarly sized models. On PinchBench, the company reports 86 percent accuracy while completing 10,000 tasks 30 percent faster than Qwen3.6 35B. NVIDIA also claims leading accuracy on the Artificial Analysis Intelligence Index, which combines nine evaluations. Those results have not been independently verified. Even so, the scope of the release goes beyond downloadable weights: data, recipes, checkpoints in different numerical formats, and support for common serving tools arrive in the same package. The model therefore supplies the compute layer of an open agent stack rather than a standalone demonstration alone.[1]
Avoiding the same model for every step
NeMo Switchyard adds an open-source routing layer that can assign an agent's work to different models for each request or each step of a multi-turn conversation. Its provider-independent kit, switchyard-libsy, considers which model can solve a task, the latency and price of the available options, and signals about system reliability. Tuning-free choices include an LLM classifier, a stage router, and an escalation router, while a prefill router can be tuned. On 145 multi-turn tasks from LangChain, NVIDIA reports that the router sent 7 percent of calls to the frontier model and cut cost by 74 percent against a frontier-only baseline. The tradeoff was a decline of about 6 accuracy points. On Cognition's FrontierCode suite, the company reports 50.6 percent at a mean cost of 3.11 dollars, placing that result 2.8 points below Opus 5 at roughly 28 percent lower mean cost. These are also NVIDIA's own measurements. The figures make clear that routing is not a cost-free gain. Lower spending comes from assigning some work to cheaper models, and the reported accuracy gap makes the price of that choice visible. The routing layer turns model selection into an operating decision made repeatedly during an agent run.[2]
Relevant guidance instead of every past lesson
ALTK-Evolve, described by IBM Research on Hugging Face, assigns a third layer to learning from past runs without changing model weights. It extracts typed guidance for strategy, recovery, and optimization, attaches support counts, and clusters near-duplicate records. IBM draws its distinction from ACE around how much material is delivered: instead of resending an entire playbook at every step, the system retrieves only the portion a model can use. In IBM's AppWorld measurements with DeepSeek-V3.2, task-goal completion reached 89.3 percent and scenario-goal completion reached 80.4 percent at 263,000 tokens per task. The reported ACE figures were 80.4 percent, 73.2 percent, and 634,000 tokens. In the gpt-oss-120b run, ALTK-Evolve used 116,000 tokens while ACE used 777,000. The measurements are the authors' own, and the post states no license. Together, the three releases describe efficiency as more than a score from one model. An open model performs the work, a router chooses which model should act and when, and a memory layer limits which lessons from earlier runs enter the context. The common constraint is that the gains reported at every layer still come from evaluations run by the organizations publishing the tools.[1], [2], [3]