DeepSeek Narrows the US-China AI Benchmark Gap to 3%
straitstimes.com

DeepSeek Narrows the US-China AI Benchmark Gap to 3%

Tech News
3 min read

Published by AINave Editorial

TL;DRBloomberg Intelligence says top Chinese models now trail US rivals by 3% on benchmark scores, down from 15% earlier in the year. DeepSeek’s LiveBench result is notable, but a narrow score gap is not the same as broad parity or a sustainable business.

DeepSeek’s September release of V4.1 Flash coincided with a sharp narrowing in the reported US-China AI benchmark gap: Bloomberg Intelligence (BI) says top Chinese models trailed US rivals by 3%, compared with about 9% in May and 15% earlier in the year. The figures describe benchmark scores, not an across-the-board measure of AI capability or national leadership.

That distinction matters. A score gap can show that leading models are converging on measured tasks, but it does not by itself settle how they perform across products, workloads or deployment conditions.

A strong LiveBench result, in a limited field

DeepSeek V4.1 Flash ranked sixth globally on LiveBench in September, the highest ranking for a Chinese model since DeepSeek’s R1 reasoning model in 2025, according to BI. The report gives the model’s last recorded LiveBench score as 81.1, compared with Anthropic’s best score of 83.4; BI analyst Robert Lea described DeepSeek’s performance as comparable to leading systems from Anthropic and OpenAI. LiveBench assesses model responses to questions, puzzles and tasks, and its rankings can change. Only three of the top 15 models were Chinese.

So the result is a meaningful sign of progress, not evidence that the wider field has evened out. The 3% cross-country benchmark comparison and DeepSeek’s sixth-place LiveBench ranking are separate measurements; neither should be read as a general capability score for every model from either country.

Domestic hardware and export controls

BI attributes China’s gains to deepening AI expertise and researchers’ ability to optimise models for domestic hardware. That progress raises questions about the effectiveness of US restrictions on technologies such as Nvidia chips, but the report does not establish that the restrictions have failed. It shows that restrictions and model progress can coexist, leaving their overall effect unresolved. The report links progress to domestic-hardware optimisation.

Better scores do not guarantee better economics

The commercial picture is less clear than the leaderboard. BI says Chinese AI firms face fierce competition, pricing pressure and difficulty monetising models, and projects that the industry could remain unprofitable until 2030. That is a conditional outlook, not a confirmed outcome: BI says a sustainable footing would depend on easing competitive pressure, industry consolidation and more rational pricing.

BI identifies ByteDance’s Doubao as the frontrunner in AI app monetisation, while DeepSeek and Tencent chatbots remain free. Those details underline why model quality and business strength should not be conflated. A model can close a benchmark gap while its maker still struggles to turn usage into durable returns. The reported monetisation and pricing picture is therefore an important counterweight to the performance gains.

FAQs

BI says top Chinese models trailed US rivals by 3% on benchmark scores after DeepSeek V4.1 Flash’s September 2026 release, down from about 9% in May and 15% earlier in the year. This is a benchmark comparison, not a measure of every dimension of AI capability. The reported figures.

Sources

Latest Tech News