One unit, two denominators

OpenAI has put GPT-5.6 Sol into a preview Ultrafast mode running on Cerebras hardware, delivering up to 750 output tokens per second. In the same week the French startup Kog demonstrated 3,000 tokens per second per request with the Laneformer 2B model, which has 2 billion parameters, on datacentre GPUs already in place. Both figures measure tokens per second.[1], [2]

Sharing a unit does not make the two figures comparable. Kog's number belongs to a model with 2 billion parameters and to a single request; OpenAI's belongs to a frontier reasoning model. The hardware differs too: Cerebras silicon on one side, installed GPUs such as AMD MI300X and Nvidia H200 on the other. Without a shared denominator, tokens per second produces no ranking.[1], [2]

Who did the measuring is the same on both sides. The Ultrafast speed figures come from OpenAI, with no independent measurement cited. Kog's 3,000 tokens per second and its claim of 30 times faster inference are the company's own, and no independent benchmark is given. Both numbers sit on the vendor-claim rung.[1], [2]

Which rung are they on?

On both sides the distance between announced speed and delivered speed is visible. Ultrafast is a preview: access runs through the API only and is limited to selected customers at first, expansion waits on capacity, and companies apply through a form. On Kog's side there is a seed round, a team of 11 people and a demonstration; the company places its Series A after running its first large model at 10 times the speed, expected in September 2026.[1], [2]

I used the same ladder earlier when reading energy capacity announcements; there too what distinguished them was the rung an announcement stood on rather than the difference in scale. That the Cerebras partnership is worth 10 billion dollars does not stand in for a speed measurement either; money does not say on what date capacity opens, or to whom.[1], [3]

What would confirm the figure?

What would put these two speeds on one scale is clear: a third-party measurement under the same model, the same precision, the same request shape and the same latency target. Kog has tied its own timetable to September 2026 and to running its first large model at 10 times the speed; OpenAI has tied Ultrafast access to capacity. If either side publishes an independent measurement by the end of September, the two figures become comparable; if neither does, tokens per second remains a vendor claim.[1], [2]