Anthropic pushes for embedded AI safety oversight

Anthropic pushes for embedded AI safety oversight

Anthropic chief Dario Amodei wants stronger independent oversight of AI. His three-stage proposal would embed third-party evaluators inside frontier developers before extending safety coordination across the industry and between governments.


Anthropic chief executive Dario Amodei has called for independent evaluators to receive continuing, employee-like access inside frontier artificial-intelligence companies as part of a wider plan to slow capability growth when safety work cannot keep pace.

The proposal would allow specialist third parties to examine not only completed AI models but also development processes, training pipelines, internal safety practices, and incidents.

Anthropic says it will adopt the embedded-evaluator model itself and wants governments to require comparable arrangements across other frontier developers.

The proposal forms the first stage of a three-part framework Amodei calls “pacing the frontier”. A second stage would establish common safety standards and limits among frontier developers in democratic countries, while a third would seek international coordination between governments.

The argument is not for an end to AI development. Amodei says model training and technical progress should continue, but at a rate that allows alignment, interpretability, security, and other safeguards to develop alongside new capabilities.

He wrote that “the stakes are too high for pacing to be an empty exercise”.

The intervention reflects growing concern that frontier systems are becoming more capable of acting autonomously before developers have mature methods for supervising every form of behaviour.

Amodei points particularly to the increasing use of AI in building subsequent AI systems and to incidents in which agentic models have undertaken actions that were not intended by their operators.

The embedded-evaluator model would change the structure of AI assurance. External testing has typically focused on assessing models at particular stages before or around deployment. Continuing access would instead give an independent party visibility into how a developer’s processes change as models are trained, updated, and used internally.

That resembles supervision in mature regulated industries more closely than conventional product testing. Amodei cites banking as a precedent for supervisors working closely with regulated organisations rather than relying entirely on periodic submissions from outside.

Artificial intelligence does not yet have an equivalent global supervisory framework. Different jurisdictions are developing different combinations of statutory regulation, voluntary commitments, model evaluations, procurement requirements, and sector-specific rules.

Permanent evaluators would also create practical governance questions of their own. Frontier developers hold commercially sensitive model information, product plans, security processes, training techniques, and intellectual property. Third parties with deep access would need clear obligations around confidentiality, competence, conflicts of interest, reporting, and escalation.

The independence of evaluators would be particularly important. A system in which companies choose, pay, and control the organisations assessing them could create doubts over whether difficult findings would be reported consistently. Government involvement could address some of those concerns but would also increase the complexity of establishing a workable regime across several countries.

Competition creates another obstacle. Frontier developers are investing heavily in computing infrastructure, researchers, data, and model training while competing for enterprise customers and technical leadership.

A company that deliberately reduces its development pace can fear surrendering ground to a rival that operates under weaker standards. Amodei’s second stage attempts to address that problem through common requirements across developers rather than relying on unilateral restraint.

International competition makes the problem still harder. Any national or regional framework has to consider what happens if developers elsewhere remain outside equivalent safeguards.

The UK already has a role in this emerging assurance landscape through its AI Security Institute and broader work on advanced-model evaluation. It is also developing sector-specific approaches in areas where AI systems interact with existing regulation.

A government-appointed healthcare AI commission recently recommended stronger lifecycle oversight and post-market monitoring, reflecting the difficulty of regulating software that can change after deployment.

Anthropic has also strengthened its international policy operation. Former UK government adviser Matt Clifford joined the company this month to lead international affairs, covering engagement outside the United States.

The embedded-evaluator proposal remains largely voluntary unless governments convert it into formal requirements. Its immediate importance lies in a major frontier developer committing itself publicly to deeper third-party scrutiny while arguing that comparable oversight should apply across its competitors.

Whether that develops into a durable regulatory model will depend on who performs the evaluations, how independence is protected, what access they receive, and whether governments can establish common standards without creating incentives for development to migrate to less restrictive jurisdictions.



  • Coalition renews push to cut electricity levies

    Coalition renews push to cut electricity levies

    More than 120 organisations want electricity levies shifted into taxation. The coalition says moving policy costs to the Exchequer could cut business electricity prices by up to 20% and support investment.


  • MPs challenge regulators over vulnerable customers

    MPs challenge regulators over vulnerable customers

    MPs say utility regulators are failing financially vulnerable customers today. The Public Accounts Committee wants better data sharing, wider social-tariff take-up, and coordinated support as energy and water debt reaches £7.2bn.


  • Ørsted gains ground in UK tax dispute

    Ørsted gains ground in UK tax dispute

    Ørsted has secured favourable guidance in a decade-long tax dispute. An arbitration opinion supports UK taxation of profits linked to Walney Extension and Hornsea 1 and could influence several related offshore wind cases.