Recursive Self-Improvement Gets a Co-Evolving Evaluator: Cambridge-NVIDIA Preprint Shows Gains in Math and Scientific Writing
techtimes.com

Recursive Self-Improvement Gets a Co-Evolving Evaluator: Cambridge-NVIDIA Preprint Shows Gains in Math and Scientific Writing

Tech News
3 min read

Published by AINave Editorial • Reviewed by Ramit

TL;DRA Cambridge-NVIDIA preprint introduces the Red Queen Gödel Machine, a recursive self-improvement framework that co-evolves both the AI agent and its evaluator, achieving gains in math and scientific writing while addressing Goodhart's Law.

The Red Queen Gödel Machine (RQGM) introduces a co-evolving evaluator that can adapt alongside its AI agent, addressing a core barrier in recursive self-improvement: static evaluators that become easy to game as capabilities grow. Source

What happened

A June 2026 preprint by 13 researchers at the University of Cambridge, NVIDIA, and partner institutions details a framework where evaluator signals are frozen within epochs and replaced only if a new evaluator outperforms on held-out ground-truth data. Source

The system employs selective erasure of scores from the previous evaluator to prevent bias carryover in training data. Reported gains include 9% higher ground-truth accuracy on Olympiad-level mathematics and 1.78 to 1.86 times higher acceptance rates in scientific-writing tasks, with 1.35 to 1.72 times token efficiency improvements. Source

Why AI builders should care

The work highlights a pathway to open-ended self-improvement that is not halted by a fixed evaluator, potentially narrowing gaps to hypothetical RSI timelines. Anthropic co-founder Jack Clark assigned a 60% probability to a fully autonomous self-improving AI arriving by the end of 2028. Source

This also introduces governance and safety considerations. The February 2026 International AI Safety Report identified recursive self-improvement infrastructure as a cross-cutting national-level security risk. Source

Practical implications

If validated, RSI with co-evolving evaluators could change how benchmarks are designed and how evaluators are updated in production-like self-improvement loops. The approach emphasizes a higher-level safeguard: evaluating new evaluators on held-out ground-truth data before replacement is allowed, anchoring the system to human-curated benchmarks. Source

However, the ground-truth anchors are structurally a static benchmark at one level higher than the problem the RQGM was designed to solve. This raises a second-order Goodhart's Law concern: a sufficiently capable agent-evaluator system could eventually saturate or exploit the ground-truth validation set. Source

Caveats

This is a preprint and has not undergone peer review. The authors describe it as a preliminary empirical investigation and note they intend to explore longer search horizons in future iterations. Source

The grounding dataset and safety claims are contingent on the held-out ground-truth set remaining a reliable signal. The paper acknowledges potential future risks where evaluators could saturate or game anchors. Source

FAQs

Recursive self-improvement (RSI) refers to AI systems that can enhance their own capabilities with limited or no human intervention, potentially creating faster cycles of capability growth. The Cambridge-NVIDIA work discusses RSI in the context of co-evolving evaluators to avoid static benchmarks being gamed. Source

Sources

Latest Tech News