Eigen RadarAI
Analysis

OpenAI adds an ultrafast tier as the AI price war stays in the middle

OpenAI opened a Cerebras tier reaching 750 output tokens per second, while Kog accelerated existing GPUs. Mid-range model discounts show that inference cost is shifting into hardware and service choices.

Artificial Intelligence··Evening
In a dark accelerator hall, blue and amber light pulses travel through distinct hardware racks and converge at one output.

Discounts in the middle, flat prices at the top

According to Ars Technica, OpenAI and Anthropic are cutting mid-range model prices as cost-conscious customers test cheaper Chinese alternatives. OpenAI cut GPT-5.6 Luna by 80 per cent, from 1 dollar to 0.20 dollars per million input tokens and from 6 dollars to 1.20 dollars per million output tokens. Anthropic launched Claude Opus 5 at 5 dollars and 25 dollars per million input and output tokens, half the price of Fable 5, and called off a September increase for Sonnet 5. Prices paid for leading US lab models have fallen by almost a quarter since mid-July on Silicon Data's token price index. DoorDash and Airbnb say they have started using Chinese-made models as both labs move some enterprise customers to usage-based billing. Headline token prices do not compare straightforwardly: Artificial Analysis found Opus 5 at medium effort matched Moonshot's Kimi K3 at max effort on performance and cost per task, while GPT-5.6 Luna at max matched DeepSeek's V4 Flash at max but cost just under twice as much per task. Mantas Lukauskas, AI tech lead at Hostinger, said prices for the very best models were flat to rising and called the changes the first real test of whether the labs can protect their top offerings.[1]

An Ultrafast tier on Cerebras hardware

According to The Decoder, OpenAI has put GPT-5.6 Sol into a preview Ultrafast mode on Cerebras hardware that reaches up to 750 output tokens per second. It sits above Fast Mode, which promises up to 2.5 times the standard speed at roughly double the price, and is open to selected customers through the API. Cerebras signed a partnership with OpenAI worth 10 billion dollars earlier in 2026. The speed figures come from OpenAI, with no independent measurement cited, and access stays limited while capacity grows; companies can register through a form. OpenAI points the mode at incident response, financial transaction monitoring, customer support, e-commerce personalisation and interactive research. While mid-tier token discounts pull headline prices down, Ultrafast turns latency and output speed into a separate paid service surface. A cheaper middle tier and a faster, capacity-limited hardware tier are sold side by side in the same model family. Which work fits which tier depends less on advertised tokens per second than on queue access and total cost per task.[2]

More inference from the GPUs already in place

According to TechCrunch, Kog, a French startup founded by Gaël Delalleau, has raised a seed round for the Kog Inference Engine, software meant to raise inference throughput on datacentre GPUs already in place, among them AMD MI300X and Nvidia H200. Varsity VC co-led the round, with support from Bpifrance, French Tech 2030 and Scaleway. In a demonstration the engine reached 3,000 tokens per second with the Laneformer 2B model, which has 2 billion parameters, and the company talks of inference 30 times faster; both figures are Kog's own and no independent benchmark is cited. Delalleau studied solid-state physics at École Polytechnique, previously founded Stribe and reached the DEFCON capture-the-flag finals four times. The team is 11 people; a May technical preview produced more than 200 business leads, the company says. Read with mid-tier discounts and the Cerebras speed tier, inference cost is no longer only the token price on a model card: hardware choice and the service layer also set the bill. Where top-tier prices stay flat, middle-tier competition pushes speed claims onto hardware. Missing independent measurement sits on both Ultrafast and Kog claims, yet inference is concentrating in the middle of the price war.[3], [1], [2]

References

  1. News sourceArs TechnicaThe price cuts land on mid-tier models while the top stays flat↩1↩2
  2. News sourceThe DecoderOpenAI opens a third speed tier for GPT-5.6 Sol on Cerebras hardware↩1↩2
  3. News sourceTechCrunchA French startup squeezes more inference speed out of today's GPUs↩