Small model, big claim: Laguna S 2.1
What stopped me in Poolside's announcement wasn't the 118-billion parameter count; it was that only about 8 billion of them run on any given token. The company has spent three years mostly selling coding models to governments and defense agencies; now it's opening the weights on Hugging Face under an OpenMDW-1.1 license, offering a no-login trial at chat.poolside.ai, and saying training took under nine weeks. In figures it supplied itself — not yet independently checked — it scores 70.2% on Terminal-Bench 2.1 and 78.5% on SWE-Bench Multilingual.[1]
I'll confess: what actually convinced me wasn't the scores, it was the architectural bet. Running roughly a twelfth of 118 billion parameters per token is a wager on “better routed” over “bigger.” The model layer is commoditizing, and value is moving up — this time it's a small lab doing the moving.[1]
Google's fast layer, slow flagship
The same day, Google shipped three models: Gemini 3.6 Flash, cutting token usage by up to 17%; Gemini 3.5 Flash-Lite, the cheapest in its class; and Gemini 3.5 Flash Cyber, a vulnerability-hunting model in a limited government pilot. But flagship Gemini 3.5 Pro hasn't moved since February; Logan Kilpatrick said the team is testing it with partners and hopes it will “land soon,” while Bloomberg has tied the delay to Google missing its own performance bar.[2]
On July 16 I wrote that the model layer was commoditizing.[3]
On July 19 I argued that the agent layer had genuinely gotten cheaper.[4]
On July 20 I assessed the Current AI development through that same lens.[5]
Today two labs teach the same lesson from opposite directions — one by staying small and routing smart, the other by holding its flagship back while rushing out the workhorse tier. My next test stays the same: whether open-weight Laguna S 2.1 gets adapted in three separate open-source projects within a month. If the gain is really architectural, it should show up there too.[1], [2]