Eigen RadarAI
Analysis

AI performance claims still lack a common yardstick

Anthropic launched Opus 5 with vendor benchmarks, while Prentis circulated an investor benchmark during an unfinished funding round. Together they show why model-performance claims still need reproducible, third-party measurement.

Artificial Intelligence··Morning
Text-free editorial illustration comparing two different AI cores with one optical apparatus

Measurements inside a release announcement

Anthropic released Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, matching Opus 4.8. Its announcement says the model leads Frontier-Bench v0.1, comes within 0.5 percent of Fable 5 on CursorBench 3.2 and triples the next-best ARC-AGI 3 score. Every result is Anthropic's own measurement; no independent verification was published.[1]

A comparison inside investor materials

Prentis, founded in April 2026, is reported to be seeking $100 million at a $1 billion valuation; the round has not closed. Investor materials claim its Hive-32B model beats GPT-5.4 and Claude Opus 4.6 on computer-use benchmarks, but disclose neither the benchmark nor the method. Both the financing stage and the model comparison therefore rest on information that remains incomplete.[2]

What would make comparison possible

The two claims do not form a ranking: one appears in a vendor release and the other in a startup's investor materials, while the models, tasks and scoring differ. Their common constraint is the absence of a reproducible method. Results from an independent party using a fixed task set, disclosed scoring and per-task consumption would turn the claims into comparable evidence. Until then, the numbers describe issuer-supplied measurements rather than a model league table.[1], [2]

References

  1. News sourceAnthropicAnthropic releases Claude Opus 5 with token pricing unchanged from Opus 4.8↩1↩2
  2. News sourceTechCrunchAI startup Prentis is in talks to raise $100 million↩1↩2