Eigen RadarAI
Analysis

Downloads and prices redraw the open-model race

Qwen’s open weights and download volume are growing as DeepSeek sharply raises prices for its new release, while OpenAI and Anthropic cut mid-tier model prices in response to cheaper Chinese rivals.

Artificial Intelligence··Night
Translucent blue model blocks spread through branching channels; a central metal scale links the flow to two stepped price paths, one coral and rising, the other teal and descending.

Small models carry distribution

Visibility and use are flowing to different places in open-weight AI. In Hugging Face’s summer review, only one repository appears in both the top 25 for likes and the top 25 for downloads. Models below 1 billion parameters account for 83 percent of all-time downloads; those above 100 billion parameters account for 1 percent. Qwen’s monthly GGUF downloads have reached 39.6 million, against 20.8 million for Gemma and 7.5 million for Llama. Its 151,448 derivatives are growing by roughly 180-210 repositories a day. Alibaba’s Qwen 3.8 family adds downloadable weights under the Apache 2.0 licence for a dense model with 27 billion parameters and the larger Qwen3.8-2.4T-A95B. The model with 27 billion parameters handles text, images, video, diagrams and documents. Its native context window is 262,000 tokens, extendable towards 1 million tokens through YaRN. Qwen says it beats the larger Qwen3.7-Plus on coding and office tasks, though no independent evaluation is provided. The figures place locally runnable and adaptable families, rather than only the largest models, at the centre of open distribution.[1], [3]

Mid-tier cuts meet a DeepSeek increase

Hosted models are moving in opposite price directions. OpenAI cut GPT-5.6 Luna from 1 dollar to 0.20 dollars per million input tokens and from 6 dollars to 1.20 dollars per million output tokens. Anthropic launched Claude Opus 5 at 5 dollars per million input tokens and 25 dollars per million output tokens, half the price of Fable 5, and cancelled a September increase for Sonnet 5. Prices paid for leading US models have fallen by almost a quarter since mid-July on Silicon Data’s index; DoorDash and Airbnb say they are using Chinese-made models to contain costs. DeepSeek chose another course with V4-Pro-0813. From 17 August, its schedule raises prices by between 50 percent and 1,100 percent according to model, token type and time of day, with separate peak and off-peak rates. V4 Flash previously cost 0.14 dollars per million input tokens and 0.28 dollars per million output tokens. Stronger agent capabilities are claimed for the new Pro build, but independent measurements have not established whether it changes the earlier performance picture. Chinese competition is therefore squeezing US mid-tier prices while DeepSeek tests whether a new release can support higher, time-sensitive charges.[2], [4]

The sticker price is only part of the cost

A simple ranking by token price would mislead. A stronger model may finish with fewer tokens or attempts, and effort settings change compute use. Artificial Analysis found Opus 5 at medium effort matched Moonshot’s Kimi K3 at maximum effort on performance and cost per task. GPT-5.6 Luna at maximum effort matched DeepSeek V4 Flash at maximum effort but cost nearly twice as much per task. Independent results for DeepSeek’s new Pro build are not yet available, so the value behind its 17 August increases remains uncertain. Open weights move the calculation beyond an API schedule to hardware, hosting, fine-tuning and operations. Qwen’s claim that its model with 27 billion parameters beats Qwen3.7-Plus is likewise unverified, but small models’ download share and derivative count show developers looking beyond launch rankings. Two distribution routes are taking shape: downloadable families compete through licences and model sizes, while closed services compete through token prices, cost per task and subscription terms. The practical choice depends on which model completes the work, how many attempts it needs and whose infrastructure runs it.[1], [2], [3], [4]

References

  1. News sourceHugging FaceSmall models carry open-model downloads, and Qwen collects the derivatives↩1↩2
  2. News sourceArs TechnicaThe price cuts land on mid-tier models while the top stays flat↩1↩2
  3. News sourceThe DecoderQwen 3.8 arrives with downloadable weights under Apache 2.0↩1↩2
  4. News sourceTech Wire AsiaWith the new build, DeepSeek raises prices by as much as 1,100 per cent↩1↩2