Eigen RadarAI
Analysis

The numbers that put new models first do not hold on the next count

Zhipu says GLM-5.3 beat Western models on one bug-finding test and trailed on others. DeepSeek's V4 Flash topped leaderboards then finished half of hard agent tasks. Hugging Face's count puts Qwen below Alibaba's 3 billion downloads.

Artificial Intelligence··Midday
In a bright evaluation hall, tall colored columns face a shorter matte set being measured by a robotic caliper arm.

GLM-5.3 wins one security test on Zhipu's numbers

Zhipu said GLM-5.3, released last week, beat Fable 5 and GPT-5.6 Sol on CyberGym, a test of finding software vulnerabilities. The figures come from the company's own announcement and have not been independently verified. Zhipu also said the model found 2,436 vulnerabilities across 269 Chinese codebases, 1,097 of them medium-to-high severity, the oldest dating back about 40 years, and that it began to form plans for complete exploitation chains rather than isolated flaws. On other security and coding tests the same model performed worse than western models. No independent evaluation of those CyberGym numbers has been published.[1]

V4 Flash finishes half the hard agent tasks

DeepSeek's V4 Flash has topped model leaderboards since its rollout. In testing by Composio it completed 53.8 percent of a batch of 30 deliberately difficult multi-step tasks built on live tools. The runs went through 8 different agent harnesses, including Claude Code, Codex and OpenCode. Those results landed on the day DeepSeek raised its prices. The company's own pricing page now lists separate peak and off-peak rates: for deepseek-v4-flash, 1 million cache-miss input tokens cost 0.44 dollars at peak and 0.22 dollars off-peak; the same line for deepseek-v4-pro costs 1.32 dollars at peak. Peak hours are defined as 01:00-04:00 and 06:00-10:00 UTC.[2]

Qwen's download count falls short of Alibaba's figure

Alibaba said its Qwen models had passed 3 billion downloads and more than 300,000 derivative models. Hugging Face's own data shows 2,061 million downloads across all repositories and 151,448 derivatives. The 2026-only count is 2,045 million. The derivative figure still runs to 2.6 times Meta's entire footprint on Hugging Face, so the ranking itself does not change. Downloads measure activity inside that one hub, leaving out API use, private deployments and Alibaba's own ModelScope platform.[3]

References

  1. News sourceThe RegisterGLM-5.3 leads on one security test and trails on the others↩
  2. News sourceVentureBeatV4 Flash tops the leaderboards, then finishes half the hard agent tasks↩
  3. News sourceThe Next WebHugging Face's own count puts Qwen below the figure Alibaba gives↩