Eigen RadarAI
Analysis

DeepSeek approaches US models on LiveBench, with a reported gap of 3 percent

Bloomberg Intelligence analyst Robert Lea puts the US–China model gap on LiveBench at about 3 percent, following DeepSeek V4.1 Flash’s September gains. The comparison gives DeepSeek 81.1 against Anthropic’s best score of 83.4. This single Bloomberg account concerns benchmark performance; only three of the top fifteen models were Chinese, and rankings can change.

Artificial Intelligence··Morning
A technician in a lavender shirt, seen from behind, examines two distinct workstation screens in a bright computing laboratory.

Lea puts the benchmark gap at about three percent

Bloomberg Intelligence, Bloomberg’s research service, has issued a new assessment comparing Chinese AI models with US-developed rivals on LiveBench. Senior analyst Robert Lea’s October 5 comparison gives a gap of around 3 percent. Earlier this year he put it at 15 percent; by May, it stood at roughly 9 percent. The Business Times carries this single Bloomberg account of the comparison, made after DeepSeek introduced V4.1 Flash in September.[1]

DeepSeek’s 81.1 is compared with Anthropic’s 83.4

LiveBench evaluates how models answer and analyse tasks, including questions and puzzles. In September’s global ranking, sixth place went to DeepSeek V4.1 Flash. The latest score cited for it is 81.1, against a best Anthropic result of 83.4. Lea regards those results as comparable with leading Anthropic and OpenAI systems.[1]

The numerical gap concerns this benchmark. It does not establish the same difference for every application, and the account does not identify the Anthropic model behind the 83.4 score.[1]

Chinese models occupy three of the top fifteen places

Chinese systems held three places in the leaderboard’s top fifteen. A model approaching the highest scores therefore sits within a broader field in which other systems remain more numerous. Lea points to growing technical expertise in China and work to adapt models to domestic hardware as factors in the gains. Rankings, he cautions, can move. His new comparison is the current development; V4.1 Flash itself was introduced in September.[1]

References

  1. News sourceThe Business Times / BloombergBloomberg Intelligence reports a narrowing US–China model gap on LiveBench↩1↩2↩3↩4