Where the cut landed
OpenAI took GPT-5.6 Luna from 1 dollar to 0.20 dollars per million input tokens and from 6 dollars to 1.20 dollars per million output tokens, a cut of 80 per cent. Anthropic launched Claude Opus 5 at 5 dollars per million input tokens and 25 dollars per million output tokens, half the price of Fable 5, and called off a September increase for Sonnet 5. Across leading US labs the prices customers actually pay have fallen by almost a quarter since mid-July on Silicon Data's token price index.[1]
The pattern matters more than any single number. Both labs cut the middle of their range and left the top alone; Mantas Lukauskas at Hostinger, which has used large language models since 2020, describes prices for the very best models as flat to rising. Both are also moving some enterprise customers off flat subscriptions and onto usage-based billing, so a team's bill now tracks the compute it consumes. The most plausible reading is that competition from freely downloadable Chinese models bites hardest in the middle of the range, where a substitute is easy to find, while a customer who needs the strongest model still has nowhere cheaper to go. Two alternatives survive the evidence: the cuts may follow ordinary product-line structure, which is what a person close to Anthropic said about Opus 5 sitting below Fable 5, or they may be an attempt to move demand toward cheaper models ahead of planned public offerings.[1]
Cost per task, at a stated effort
When Opus 5 launched in July I argued in this column that the operative change for a developer was a per-request effort dial, and that every measurement supporting it came from the vendor, which left the dial untested. Artificial Analysis has now supplied the outside comparison that was missing: Opus 5 at medium effort matched Moonshot's Kimi K3 at max effort on both performance and cost per task, and GPT-5.6 Luna at max matched DeepSeek's V4 Flash at max while costing just under twice as much per task. The unit in that comparison is a completed task at a named effort level, which is the number a team can hold against its own workload.[1], [3]
What the comparison leaves out is the rest of the workflow. Cost per task on a benchmark suite counts neither the retries a weaker model provokes, nor the reviewer time a wrong answer consumes, nor the rollback when an agent commits a bad change. Silicon Data's index measures prices paid, which moves with the mix of models customers buy as well as with list prices, so a decline of almost a quarter can partly reflect teams stepping down a tier rather than any single price falling. A team that wants a real answer has to run its own tasks at two effort settings and record failures alongside spend.[1]
What makes a swap practical
The reporting behind those cuts names freely downloadable Chinese models as part of the pressure, and a concrete instance of what a team can download arrived this week. Alibaba's Qwen team released Qwen 3.8 under the Apache 2.0 licence, with a multimodal dense model carrying 27 billion parameters and a larger Qwen3.8-2.4T-A95B variant, weights on Hugging Face and ModelScope, a native context window of 262,000 tokens and a thinking mode that can be turned off per query. A permissive licence and downloadable weights are what let a team move a workload without reopening a contract; the claim that the model with 27 billion parameters beats Qwen3.7-Plus on coding and office tasks is the team's own, and the release names no independent evaluation. Downloadable weights also carry their own bill in serving hardware and operations, which is where a saving on list price can quietly go.[2], [1]
So the honest test of how far this goes is narrow and datable. If Silicon Data's index still shows mid-range prices from the two US labs falling while their top-tier prices hold flat or higher at the end of December 2026, the substitution these cuts enable stays confined to the middle of each line-up, and the question of whether the model layer is becoming interchangeable will have been answered only for the mid-range. The signal to watch is a published cost-per-task comparison in which an open-weight model at max effort matches a US top-tier model at medium, on tasks a team actually runs.[1], [2]