US labs cut mid-tier prices as open Chinese models fill the catalogue
As OpenAI and Anthropic cut mid-tier API prices, Hugging Face, Qwen and Meta open weights give buyers cheaper endpoints and a commercially usable model layer at once.
Artificial Intelligence··Night
Mid-tier API cuts and enterprise bill pressure
OpenAI cut the mid-tier GPT-5.6 Luna model from $1 to $0.20 per million input tokens and from $6 to $1.20 per million output tokens. Anthropic opened Claude Opus 5 at $5 and $25 per million tokens, half the published price of its Fable 5 model, and this week cancelled a September rise planned for Sonnet 5. Silicon Data's token price index shows the price paid for leading United States models down by almost a quarter since mid-July; the cuts cover mid-tier products while flagships stay outside the discount. Artificial Analysis found Opus 5 at medium effort close to Moonshot's Kimi K3 at max effort on cost per task, while GPT-5.6 Luna at max effort matched DeepSeek's V4 Flash but cost just under twice as much per task. DoorDash and Airbnb said they started using Chinese-made models to rein in bills. Anthropic and OpenAI declined to comment; a person close to Anthropic said Opus 5 sitting below Fable 5 follows from how the model family is built and has no connection to competitors.[1]
The Hugging Face catalogue and Qwen3.8-27B
Hugging Face compiled open-model activity on its own platform for the first seven months of 2026. In almost every month, the largest open model from a Chinese lab was bigger than anything a United States lab shipped: Chinese releases ranged from 754 billion to 2.78 trillion parameters, against under 130 billion parameters in most months on the United States side. Of 178 Chinese releases above 20 billion parameters, 59 per cent carry Apache 2.0 and 22 per cent carry MIT, and none restricts commercial use; United States labs left the licence undeclared on 30 per cent of theirs. Attention gathers narrowly: 85.6 per cent of models stay under 200 lifetime downloads, and 1.5 per cent of repositories collect 99.2 per cent of all downloads. Qwen's derivative count has reached 151,448. In July 2026 Claude Code took 44.4 per cent of agent traffic and Codex 20.8 per cent. The compilation measures only Hugging Face activity. Beside that Alibaba's Qwen team released Qwen3.8-27B under Apache 2.0 with a native context window of 262,144 tokens, extendable to 1,000,000 with YaRN. Qwen lists 61.7 on SWE-bench Pro, 89.2 on GPQA Diamond and 84.3 on OSWorld-Verified; those are vendor figures without independent measurement.[2], [3]
A local agent model and the shared price–open-weight picture
Meta put Muse Glimmer under Apache 2.0 with weights on Hugging Face: a 30 billion parameter model for agent loops, tool calls and coding that run locally on a workstation. Meta recommends 24 to 32 GB of unified memory or VRAM; a 4-bit quantised copy needs roughly 17 to 20 GB. The company reports strong class results on SWE-Bench, DeepSearch QA, τ-Bench and MCP-Atlas, and claims better multi-step tool reliability than Gemma 4 31B and Qwen 3.6 27B. Those comparisons are Meta's own, without independent verification. Meta also cites up to a 3.1 times throughput gain from DFlash speculative decoding on Apple Silicon and NVIDIA RTX cards. Weights ship under a permissive licence; training data and pipeline stay outside it. Taken with the mid-tier API cuts, the Hugging Face catalogue and the Qwen open-weight drop, buyers face cheaper United States mid-tier endpoints while a commercially usable open layer, including a local agent stack, expands under permissive licences. How far those open scores hold under independent tests remains open, and download concentration still funnels attention into a thin repository slice.[4], [1], [2], [3]