
DeepSeek Narrows the US-China AI Benchmark Gap to 3%
Published by AINave Editorial
DeepSeek’s September release of V4.1 Flash coincided with a sharp narrowing in the reported US-China AI benchmark gap: Bloomberg Intelligence (BI) says top Chinese models trailed US rivals by 3%, compared with about 9% in May and 15% earlier in the year. The figures describe benchmark scores, not an across-the-board measure of AI capability or national leadership.
That distinction matters. A score gap can show that leading models are converging on measured tasks, but it does not by itself settle how they perform across products, workloads or deployment conditions.
A strong LiveBench result, in a limited field
DeepSeek V4.1 Flash ranked sixth globally on LiveBench in September, the highest ranking for a Chinese model since DeepSeek’s R1 reasoning model in 2025, according to BI. The report gives the model’s last recorded LiveBench score as 81.1, compared with Anthropic’s best score of 83.4; BI analyst Robert Lea described DeepSeek’s performance as comparable to leading systems from Anthropic and OpenAI. LiveBench assesses model responses to questions, puzzles and tasks, and its rankings can change. Only three of the top 15 models were Chinese.
So the result is a meaningful sign of progress, not evidence that the wider field has evened out. The 3% cross-country benchmark comparison and DeepSeek’s sixth-place LiveBench ranking are separate measurements; neither should be read as a general capability score for every model from either country.
Domestic hardware and export controls
BI attributes China’s gains to deepening AI expertise and researchers’ ability to optimise models for domestic hardware. That progress raises questions about the effectiveness of US restrictions on technologies such as Nvidia chips, but the report does not establish that the restrictions have failed. It shows that restrictions and model progress can coexist, leaving their overall effect unresolved. The report links progress to domestic-hardware optimisation.
Better scores do not guarantee better economics
The commercial picture is less clear than the leaderboard. BI says Chinese AI firms face fierce competition, pricing pressure and difficulty monetising models, and projects that the industry could remain unprofitable until 2030. That is a conditional outlook, not a confirmed outcome: BI says a sustainable footing would depend on easing competitive pressure, industry consolidation and more rational pricing.
BI identifies ByteDance’s Doubao as the frontrunner in AI app monetisation, while DeepSeek and Tencent chatbots remain free. Those details underline why model quality and business strength should not be conflated. A model can close a benchmark gap while its maker still struggles to turn usage into durable returns. The reported monetisation and pricing picture is therefore an important counterweight to the performance gains.






















