Eigen RadarAI
Analysis

AllSpark releases the open weights of Iris-mini and Iris-pro with the harness that ran their benchmarks

The AI lab AllSpark has opened up Iris-mini and Iris-pro, two search agents that carry out web research on their own, by releasing their weights. The smaller model holds 35 billion parameters and the larger 397 billion; both are built from Qwen-series bases with a 256,000-token context window. The team says each leads its size class among open-weight search agents, and it publishes the agent loop, tools and evaluation code behind those scores as the Iris Harness, which connects to any OpenAI-style endpoint.

Artificial Intelligence··Night
A technician examines a circuit board on a test bench linked by colored cables to an open grey equipment cabinet in a bright hall.

Weights on Hugging Face, code and harness on GitHub

The AI lab AllSpark has placed the Iris-mini and Iris-pro weights in a collection on Hugging Face, the model-sharing platform, and the code in a repository on GitHub. The release goes beyond the models: the Iris Harness carries the agent loop that runs them, the tools they call, the context-management strategies and the four benchmarks used in the measurements, complete with evaluation code. The harness connects to any model server that speaks OpenAI's interface format.[1]

Context management alone made a difference of as much as 21.2 BrowseComp points

The team ran every benchmark twice, once with runtime context management switched on and once with it off, leaving the tools, the context limits and the judge model unchanged between the pair. The difference is large: on BrowseComp, a benchmark that measures an agent's ability to track down hard-to-find answers on the web, Iris-mini scored as much as 21.2 points higher with context management on. That gap exceeds the differences usually reported between systems.[1]

Two sizes on a shared Qwen base, and the team's reported class lead

Iris-mini has 35 billion parameters; Iris-pro goes up to 397 billion. Both are built from open models of the Qwen series and read a 256,000-token context window, a token being the unit of text a model processes at a time. The team says each model leads the open-weight search agents of its size class: the BrowseComp score is 82.2 for Iris-mini and 88.6 for Iris-pro. Those scores are the team's own measurement; no independent measurement exists yet.[1]

References

  1. News sourceThe DecoderAllSpark opens the weights of Iris-mini and Iris-pro and ships the agent loop that ran the benchmarks↩1↩2↩3