Google, DeepSeek and OpenAI turn model speed into a pricing race
Gemini 3.7 Flash, DeepSeek’s peak-hour tariff, OpenAI’s Ultrafast mode and Writer’s rebuilt harness shift model competition from benchmark scores toward delivery speed and operating cost.
Artificial Intelligence··Evening
Gemini 3.7 Flash: introductory price and vendor scores
Google has released Gemini 3.7 Flash, which the company calls its most intelligent workhorse model yet for coding and agents. On the company’s own figures the model scores 43.6 per cent on FrontierCode 1.1 Main against 34.4 per cent for 3.6 Flash, 65.3 per cent on DeepSWE v1.1 against 49.0 per cent, and an Elo of 1588 on WebDev Arena against 1538. These are vendor measurements published with the launch, not independent lab results. Introductory rates run at 0.75 dollars per million input tokens and 3.75 dollars per million output tokens through 31 December 2026, rising afterwards to the published 1.50 dollars and 7.50 dollars. Availability covers Google Antigravity, the Gemini API in Google AI Studio and Android Studio, the Gemini Enterprise Agent Platform, and Gemini Spark for Google AI Pro and Ultra subscribers in more than 160 countries.[1]
DeepSeek’s hourly tariff and OpenAI’s Ultrafast mode
At DeepSeek the price of inference now depends on when the request arrives. From 16 August 2026 the company charges 3.96 dollars per million output tokens for V4 Pro at peak hours and 1.98 dollars off-peak, against 0.87 dollars before. V4 Flash moves from 0.28 dollars to 1.32 dollars at peak and 0.66 dollars off-peak, with off-peak rates at half the peak level for both models. DeepSeek said the structure lets it allocate resources more reasonably; input token rates were not published separately. Earlier promotional rates had been due to expire on 31 May, and the company had previously said it would make them permanent before reversing that decision. In the same window OpenAI opened a limited preview of Ultrafast, a mode it says runs GPT-5.6 Sol at 14 times standard speed and delivers up to 750 output tokens per second. Access goes for now to a small group of customers, and the company links the speed to its partnership with the chipmaker Cerebras. Both figures are OpenAI's own, published without an independent measurement or a stated workload, and pricing was not disclosed. Access is to widen as capacity grows.[2], [3]
Writer’s rebuilt harness and the shared cost pressure
Writer released Palmyra X6 on Thursday alongside a rebuilt agentic harness. Palmyra X6 is a post-training variation on Z.ai’s open-source GLM-5.2, and the harness work targets multi-step tasks that burn tokens without finishing. The company puts the average cost reduction from the harness changes alone at 40 per cent, and says basic tasks can run up to 50 per cent cheaper. Writer’s researchers write that harness efficiency changes proved more reliable than model selection for cutting cost, and that the harness multiplies efficiency across every model an organisation runs. Chief executive May Habib said enterprises are sick of chasing the next benchmark and want flattening cost. The model reached Writer’s clients on Thursday; every figure is the company’s own and no independent evaluation has been published. Google’s time-boxed tariff, DeepSeek’s peak-hour split, OpenAI’s speed mode and Writer’s harness-centred cost claim land in the same news window, moving the competitive centre of gravity from intelligence tables toward delivery speed and operating cost.[4], [1], [2], [3]