The company's own table
DeepSeek said that in V4-Flash-0731, now in public beta, it changed neither the architecture nor the parameter count, and that the build is a training and post-training upgrade. In the figures it published itself, Terminal Bench 2.1 rises from 61.8 to 82.7 against the preview build, NL2Repo from 39.4 to 54.2, Cybergym from 38.7 to 76.7 and DeepSWE from 7.3 to 54.4. In the same table Claude Opus 4.8 stands at 85.0 on Terminal Bench 2.1, and on DSBench-FullStack Opus 4.8 scores 71.6 against 68.7 for DeepSeek's new build. All of these are the company's own measurements, and no independent evaluation exists yet.[1]
For a developer the honest comparison is not a rival model but the preview build of the same weight family. On that axis three of the four numbers behave alike, with gains between fifteen and thirty-eight points. The move from 7.3 to 54.4 on DeepSWE sits outside that pattern and carries forty-seven points on its own. On the same architecture, the same parameter count and the same benchmark, a difference of that width is not self-explained by the one variable the announcement names, which is the post-training work.[1]
The second thing that changed in the same release
The same release began supporting the Responses format natively and was adapted for Codex. On agentic benchmarks the score does not come only from the text a model produces; whether the layer carrying tool calls can format a call correctly and return the answer also enters the score. That makes it a plausible reading that part of the width on DeepSWE comes from a better fit with the measuring apparatus rather than from model competence. The competing account is equally strong: the preview build may simply have been undertrained for tool use, in which case the forty-seven points are a real difference in capability.[1]
The experiment that separates the two readings is not expensive. Run the two builds side by side while holding the tool layer, the task set and the scoring method constant, and it becomes directly visible how much of the difference comes from the apparatus. The table the company published does not say which of those conditions were held fixed. A benchmark score is not a workflow result, and until the contributing layer is identified the number does not carry information a purchasing decision can rest on.[1]
The number underneath the price
Pricing is stated as $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit and $0.28 per million output tokens. The comparison figures given are $5 and $30 for GPT-5.6 Sol, $10 and $50 for Claude Fable 5 and $3 and $15 for Claude Sonnet 5. That range is what makes a team's retry budget, wide fan-out and a second review pass affordable enough to plan around.[1]
Price per token and cost per completed task are not the same thing: when a cheap model needs more attempts, more review and more error recovery, the gap closes. The signal that would measure that distinction sits in an independent evaluation. If, by 31 October 2026, an independent measurement with a disclosed tool layer reproduces the Terminal Bench 2.1 result within a few points, the gain is on the model side; if the result comes in markedly lower, the apparatus-fit account gains weight.[1]