What does a token count?
OpenRouter analyst Peter Walker put a date on the crossover: after 6 February 2026, agent token use on the marketplace went from 0.51 trillion to 7.3 trillion, a 14-fold rise, while human use grew 2.8-fold over the same stretch. Two figures on one platform, over one period, on one denominator: a rare piece of AI demand data that arrives already normalised.[1]
Then read what sits inside the count. About 70 percent of that agent traffic is cached prompts, billed at lower rates, and OpenRouter's traffic leans towards open-weight models that spend more tokens per task than OpenAI's or Anthropic's. So the 14-fold figure counts billed marketplace tokens whose composition changed while they were being counted, and it tracks how the work is priced at least as much as how much compute it took. The alternative reading also stands: demand may really have risen 14-fold, with caching passing an efficiency gain to buyers. The same figures fit both, and that is the difficulty.[1]
Two prices from the same weekend
DeepSeek moved its own price over the same weekend. From 00:00 Beijing time on 23 August the company drops the peak and off-peak split on Saturdays and Sundays and charges the off-peak rate for both full days. Off-peak is half of peak: on the flagship DeepSeek-V4-Pro, output goes from 27 yuan to 13.5 yuan per million tokens. A provider selling two of seven days at half price is telling you where its serving capacity is tight and where it sits idle.[2]
The other price went the other way. Bloomberg's report, carried by the Verge, is that some of Nvidia's biggest customers, the firms that build the AI data centres for companies such as Oracle and Microsoft, have been told server prices are rising more than 15 percent, on top of GPU price rises already taken this year. That is the cost of the machine underneath the token, and it went up in the same days the price of a weekend token came down.[3]
Which number gets tested, and when?
Put the two side by side and the mechanism looks ordinary. Serving capacity became more expensive to build, and providers ration and reprice it, through off-peak bands on one side and cached-prompt discounts on the other. Here is the test: if DeepSeek keeps the all-weekend band past 31 August and Nvidia's quarterly results address server pricing, supply cost is what moved the price and a temporary promotion can be ruled out. Watch DeepSeek's published API price table on the first weekend of September, and the server-price line in Nvidia's quarterly materials.[2], [3]
One more thing about that 14. On 16 August this column found two tokens-per-second figures that could not share a denominator, and the OpenRouter series shows the same gap one layer up, where the token carries a price rather than a measurement. A token is a billing line. Serving capacity is measured in racks, watts and hours, and none of those appear inside a 14-fold count. Until a marketplace publishes cached and uncached tokens as separate series, the honest reading of the number is that agent traffic grew a lot and that we cannot yet say how much silicon it took.[1], [4]