New model and service choices split price, speed and capability
Google opens Gemini 3.7 Flash with introductory rates and vendor scores while Cerebras pitches OpenAI Ultrafast at up to 750 tokens per second. DeepSeek, separately, raises interface prices from 16 August.
Artificial Intelligence··Night
Gemini 3.7 Flash: introductory price and vendor scores
Google has made Gemini 3.7 Flash generally available for coding and agent workflows. Through 31 December 2026 the introductory rate is 0.75 dollars per million input tokens and 3.50 dollars per million output tokens; from 1 January 2027 those figures rise to 1.50 dollars and 7.50 dollars per million tokens. The model is available to developers through the Gemini API in Google AI Studio and Android Studio, to enterprises through the Gemini Enterprise Agent Platform and the Gemini Enterprise app, and to consumers through Gemini Spark for Google AI Pro and Ultra subscribers in more than 160 countries. On benchmark results Google published itself, the model scores 43.6 per cent on FrontierCode 1.1 and 65.3 per cent on DeepSWE v1.1, against 34.4 per cent and 49.0 per cent quoted for 3.6 Flash. Those comparisons rest on the vendor’s own account and carry no independent verification. Inside one product family, then, a time-bounded price window and a vendor claim on coding and software-engineering tests arrive together, so a buyer has to weigh the post-introductory rate against scores that outsiders have not yet checked.[1]
Ultrafast on Cerebras hardware: speed as its own tier
Cerebras and OpenAI have introduced a service tier called Ultrafast that runs GPT-5.6 Sol at up to 750 output tokens per second. For now it is a preview open only to selected customers; nothing was said about when access will widen or what the tier will cost. On Cerebras’s own account the tier runs on the company’s wafer-scale processor architecture, where 44 GB of SRAM per chip keeps model weights on the chip itself. The company also reports that Humanity’s Last Exam was completed in 11 hours 11 minutes and that the GDP-Val measurement showed an end-to-end speedup of 5.6 times. All of these comparisons rest on measurements by Cerebras and OpenAI themselves, with no independent verification published. Speed is offered here as a separate product surface: output rate and exam duration sit in the foreground, while price and broader access are simply absent from the announcement. What a reader can hold is that the capability claim travels in a tier that is not yet tied to a public price list or to an independent scorecard of model quality.[3]
DeepSeek’s price schedule and three separate axes
DeepSeek’s new interface price schedule takes effect at 16:00 UTC on 16 August. On the schedule The Decoder compiled, V4-Pro cache-hit rates are 0.003625 dollars per million tokens on the old line, 0.022 dollars off-peak and 0.044 dollars at peak; cache-miss input is 0.435 dollars on the old line, 0.66 dollars off-peak and 1.32 dollars at peak, while output is 0.87 dollars, 1.98 dollars and 3.96 dollars on the same steps. Peak hours are defined as 01:00–04:00 and 06:00–10:00 UTC, which line up with the Chinese working day. The same announcement took V4-Pro build 0813 out of its testing phase; on figures the company supplied, its Terminal Bench 2.1 result went from 72.1 to 87.9 and its DeepSWE result from 12.8 to 62.7. Google’s side shows an introductory price window and vendor scores, the Cerebras–OpenAI line shows output tokens per second, and DeepSeek shows a cache- and peak-hour-tiered price rise. The three developments do not form one operator chain; read together, they show model and service choices presenting price, speed and capability on separate surfaces rather than in a single bundled offer.[2], [1], [3]
Related columns
For more information on this topic, you can read the related columns.