Same price list, new decision point
Anthropic released Claude Opus 5 on July 24 at $5 per million input tokens and $25 per million output tokens — the same price as Opus 4.8. The company says it comes close to the capability of its own top model, Claude Fable 5, at half the price, and it is reachable through Claude.ai, the API, Claude Code and Claude Cowork.[1]
The same announcement adds three control surfaces: an effort setting, mid-conversation tool changes, and an automatic fallback that routes a request to another model when a safety classifier fires. The last two are in beta. To my mind that is where the change for a software team actually sits: the choice becomes how much effort to buy for a single request rather than which model to call. The opposite reading is available too — the effort setting may be a new name for a sampling budget that has existed for a long time, changing nothing in the workflow.[1]
Every measurement comes from the vendor
The numbers in the announcement are ambitious: leading other models on Frontier-Bench v0.1 with twice the Opus 4.8 result, landing within 0.5 percent of Fable 5 on CursorBench 3.2, and scoring three times as high as the next-best model on ARC-AGI 3. All of it is Anthropic's own measurement, and the announcement points to no independent verification.[1]
The version number deserves attention: Frontier-Bench v0.1 is a benchmark set in its first release. At that stage it is hard to separate a model's progress from the benchmark's own construction, because no second party has yet run the same task set with the same scoring. The cost claims carry the same problem: phrases such as half the cost and one-third the cost fold together the unit price and the number of tokens a task consumes, and neither is given separately. The results may well hold up in independent runs, and if they do this objection falls away.[1]
Where the control went
One condition would make the dial measurable: if by October 31, 2026 a party other than Anthropic publishes Frontier-Bench v0.1 or CursorBench 3.2 results holding the task set and scoring method constant across Opus 5 and at least one non-Anthropic model, the cost difference the effort setting promises becomes a testable number. If no such publication appears, what remains should be read as a vendor comparison.[1]
Meanwhile the practical cost of the automatic fallback is already visible. Routing a request to another model when the classifier fires removes the error message but adds the question of which model produced the answer. For a team that needs reproducibility the consequence is concrete: logging which model served each request stops being an optional detail.[1]