Qwen’s download race runs into a measurement problem
Alibaba touts three billion downloads while Hugging Face’s own count tells a smaller story; user-specific benchmarking tools show why a single total cannot settle the model race.
Artificial Intelligence··Evening
Alibaba puts three billion downloads forward
Alibaba’s Qwen models reached 3 billion downloads over the past six months, according to the Hugging Face open-model report relayed by Fortune. The same report puts Google at 418 million downloads and Meta at 227 million. Fortune’s Bloomberg-bylined report says the Qwen family open-sourced more than 460 models and that its ecosystem has spawned more than 300,000 derivative models. A download total shows which models developers choose to fine-tune and deploy; it does not show how many people use those weights in a finished product. The report says the figures come from Hugging Face’s state-of-open-models report, dated August 14. That report places Qwen’s download volume alongside Google and Meta on the same comparative scale. Alibaba offered a broad open-model family during this period, and the download total is read as a measure of how fast that family moved across Hugging Face. The report supplies no separate figure for end-user reach. Fortune separates the point clearly: the total reflects fine-tuning and deployment choices, not finished-product use. On first reading, the 3 billion figure draws a clean leadership table in the open-model race; what the table counts becomes the next measurement problem.[1]
Hugging Face’s own count draws a smaller table
The Next Web reports that Hugging Face’s own count puts Qwen below the figure Alibaba gives. Alibaba said its Qwen models had passed 3 billion downloads and more than 300,000 derivative models. The platform’s own data shows 2,061 million downloads across all repositories and 151,448 derivatives. The 2026-only count is 2,045 million. The derivative figure still runs to 2.6 times Meta’s entire footprint on Hugging Face, so the ranking itself does not change. Downloads measure activity inside that one hub. API use, private deployments and Alibaba’s own ModelScope platform sit outside the count. Two download totals now stand side by side for the same story: the company announcement says 3 billion, while the platform count gives 2,061 million downloads and 151,448 derivatives. The Next Web’s account shows that, even when the ranking holds, how the absolute figure is calculated still matters to the reader. When the scope, period and repositories change, the leadership table looks different even if the unit is still called a download.[2]
User-specific benchmarking shows one total is not enough
Artificial Analysis opened a tool called Optima, The Decoder reports, so users can try models against their own data. The tool accepts existing evaluation sets, agent traces from tools such as Arize or Langfuse, or just a description of a use case. It compares models on cost per task and time per task alongside answer quality. Two scoring methods are offered: rubric-based evaluation, and putting two answers side by side and asking which one is preferred. Billing runs on the actual token costs of the models used with no markup added; it charges 0.125 dollars per criterion per model and 0.375 dollars per pairwise comparison. Early users built benchmarks for finance and accounting agents and tested matching a legal writing style. The tool is available now. In the Qwen race, figures such as 3 billion and 2,061 million measure distribution inside a hub, while Optima asks each user which model works on their own workload. The open-model contest does not close on a single download total; which repository is counted, which period is taken and which task is tested all change the result. Taken together, the three developments show that model leadership still turns on the choice of measurement more than on one headline number.[3], [1], [2]