AI performance claims still lack a common yardstick
Anthropic launched Opus 5 with vendor benchmarks, while Prentis circulated an investor benchmark during an unfinished funding round. Together they show why model-performance claims still need reproducible, third-party measurement.
Artificial Intelligence··Morning
Measurements inside a release announcement
Anthropic released Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8. Its announcement says the model leads Frontier-Bench v0.1, comes within 0.5 percent of Fable 5 on CursorBench 3.2 and triples the next-best ARC-AGI 3 score. Every result is Anthropic's own measurement; no independent verification was published.[1]
A comparison inside investor materials
Prentis, founded in April 2026, is reported to be seeking $100 million at a $1 billion valuation; the round has not closed. Investor materials claim its Hive-32B model beats GPT-5.4 and Claude Opus 4.6 on computer-use benchmarks, but disclose neither the benchmark nor the method. Both the financing stage and the model comparison therefore rest on information that remains incomplete.[2]
What would make comparison possible
The two claims do not form a ranking: one appears in a vendor release and the other in a startup's investor materials, while the models, tasks and scoring differ. Their common constraint is the absence of a reproducible method. Results from an independent party using a fixed task set, disclosed scoring and per-task consumption would turn the claims into comparable evidence. Until then, the numbers describe issuer-supplied measurements rather than a model league table.[1], [2]
Related columns
For more information on this topic, you can read the related columns.