Eigen RadarAI
Analysis

The model race runs from price tags to usage accounts

New code models, Chinese price pressure, Nvidia's scale ambitions and GitHub's token report show how comparisons in AI purchasing are becoming more granular.

Artificial Intelligence··Evening
Compute modules of different sizes connect to separated usage streams in front of a data center.

A cheaper release faces an even cheaper rival

Microsoft says MAI-Code-1.1-Flash, now arriving in GitHub Copilot, is 25 per cent more token-efficient than its June predecessor and costs one quarter as much; it also says developers accept 4 per cent more of the model's output. The model-card table reported by The Decoder carries the comparison beyond Microsoft's own family. On Terminal Bench 2.1, the new Microsoft model scores 62.9 per cent, versus 51.7 per cent for its predecessor and 82.7 per cent for DeepSeek-V4-Flash-0731. On SWE-bench Verified, MAI-Code-1.1-Flash reaches 72.6 per cent and its predecessor 71.6 per cent. At listed prices per million tokens, DeepSeek-V4-Flash costs 0.14 dollars for input and 0.28 dollars for output, while MAI Code 1.1 Flash costs 0.20 dollars and 1.20 dollars. Microsoft's announcement does not directly compare models outside its family. The performance and price figures also come from values published by Microsoft and DeepSeek, rather than an independent evaluation. Even so, the table lets buyers look beyond whether a release is cheaper than its predecessor and ask what a rival offers for similar work.[1]

Price pressure works in both directions

The AFP report published by The Star shows that this comparison extends beyond one code model. It says cost-efficient Chinese models are pushing Silicon Valley laboratories to lower prices as users move between American and Chinese options. OpenAI has cut fees for its lightweight Luna model by 80 per cent, while Anthropic offers a model approaching its most powerful system at half the price. DeepSeek, by contrast, has announced a significant planned increase in prices for programmers even as V4 Flash leads OpenRouter's usage ranking. At AGI Bar in Beijing's Zhongguancun technology district, the owner runs V4 Flash on local Nvidia workstations and gives customers free DeepSeek access, paying the startup nothing for that use. Analyst Jack Gold tells AFP that older, smaller or more modest models can accomplish plenty, particularly in agent tasks. This picture turns price competition into something more complex than a steady march downward. Hosting, the capability a task actually requires and a provider's price change all enter the same purchase decision; the most expensive frontier model no longer serves automatically as the starting point for every job.[2]

As scale grows, the accounting gains more lines

Meanwhile, Nvidia's expansion as a model producer shows that the scale race continues. Citing The Information, The Decoder reports that the company is working on the open-weight Nemotron 4 family, whose largest member would carry at least 1 trillion parameters. That is twice the size of Nemotron 3 Ultra; the report puts Nvidia's cloud spending to train its own models at 28 billion dollars through 2031 and says the earliest release would be this autumn. Nvidia has not issued its own Nemotron 4 announcement. On the buyer's side, GitHub has added a per-model token breakdown to the Copilot usage report: input, output, cache-read and cache-write tokens now appear beside the credits consumed by each model. The detailed breakdown appears only in a report downloaded from the AI usage page in billing settings. As producers enlarge model parameters and training expenditure, customers are beginning to separate which model and token category appears on a bill. Two scales of competition become visible at once: the large capacity a supplier builds and the individual units of consumption a user wants to track and reduce.[3], [4]

References

  1. News sourceThe DecoderMicrosoft's MAI-Code-1.1-Flash trails DeepSeek's cheaper model on the benchmark table Microsoft itself published↩
  2. News sourceThe StarAFP finds Chinese models pushing US labs to cut prices while Chinese labs plan increases↩
  3. News sourceThe DecoderNvidia is preparing a Nemotron 4 family whose largest model would carry at least 1 trillion parameters↩
  4. News sourceGitHub ChangelogToken spend is now broken down by model in the GitHub Copilot usage report↩