
Embedded AI evaluators gain momentum inside OpenAI and Anthropic as Geoffrey Hinton backs transparency demands
Published by AINave Editorial • Reviewed by Ramit
The push for embedded AI evaluators inside OpenAI and Anthropic is gaining formal structure. The AI Evaluator Forum, a coalition that includes the nonprofit evaluation group METR and the AI Verification and Evaluation Research Institute (AVERI), released a letter on Friday outlining minimum conditions for the third-party evaluators that both labs have already pledged to welcome. Signed by AI pioneers Geoffrey Hinton and Stuart Russell, the letter is the first concrete governance proposal for how external oversight should operate inside frontier AI companies.
The minimum conditions for embedded evaluators
The letter demands that evaluators be granted "scientific objectivity, transparency, independence, and robust protections against interference from the evaluated companies." More specifically, it calls for unfiltered communication with company boards and other privileged oversight bodies, the ability to release findings publicly, and access to the same systems, data, tools, and physical spaces as internal assessors. Evaluators must also be shielded from retaliation.
The letter follows commitments from Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman to welcome third-party evaluators with employee-level access. Microsoft CEO Satya Nadella also expressed support for the idea.
What changes for builders and operators
For developers, founders, and product teams building on top of OpenAI and Anthropic models, this shift toward formalized internal oversight has several practical implications.
First, it could lead to more transparent reporting of model risks. If evaluators can publish findings publicly, teams building agents or deploying models in sensitive domains may get earlier warnings about capabilities or failure modes that internal safety teams might downplay.
Second, the access requirements mean evaluators will need security clearances and formal contracts defining their scope. That adds governance overhead for the labs, but it also creates a precedent for how risk assessments are conducted. If the model you depend on is suddenly flagged for a novel risk, the process behind that flag will be more auditable.
Third, the letter explicitly calls for diverse viewpoints on risk domains, including biological safety and loss of control. That suggests evaluators will cover a broader range of scenarios than traditional red-teaming, potentially affecting which model capabilities get restricted or delayed.
The open questions and constraints
The letter is a demand, not a binding agreement. Implementation depends on how OpenAI and Anthropic respond to the specifics. The Financial Times has reported that the safety push is already creating tensions inside both companies, highlighting the challenge of turning public commitments into operational reality.
Details like evaluation frequency, funding independence, and how confidential findings interact with trade secrets remain unspecified. The plan also relies on the labs voluntarily granting the access and protections described. Without regulatory mandate, enforcement is limited to public pressure.
For now, the letter signals that the era of purely internal AI risk assessment is ending. Builders should watch how these conditions are adopted, because they will shape how model risks are communicated, and ultimately how safe or restricted the models you depend on become.
FAQs
Sources
- The 'Godfather of AI' backs a new watchdog plan to track OpenAI and Anthropic's AI risks from the inside
- Anthropic, OpenAI proposed new AI watchdogs: Why idea should...
- ‘Godfather of AI’ says tech regulation is nearing... | The Guardian
- AI bosses’ safety push sparks rift inside OpenAI and Anthropic
- AI kill switch won't work in the long run: 'Godfather' of AI - Yo...
- After OpenAI, Anthropic incidents, AI godfather warns humans may no longer control AI
- The ‘Godfather of AI’ backs a new watchdog plan to track OpenAI...
- OpenAI & Anthropic Insider Quits With a Terrifying Warning - YouTube
- Canadian experts back call for independent watchdogs at the world’s top AI companies
- 'Godfather of AI' Says US Has About a Year to Regulate AI - Business Insider




















