Anthropic chief executive Dario Amodei has called for the pace of frontier artificial intelligence development to be slowed, warning that rapid advances in AI capabilities could soon outstrip efforts to understand, evaluate and control increasingly powerful systems.
Amodei, in an essay titled We Must Pace the Frontier, argued that the industry should not halt AI development but should give safety and alignment work enough time to keep pace with advances in model capabilities.
He said the need for a more measured approach had become more urgent in recent months as AI systems have increasingly demonstrated the ability to contribute to the development of subsequent generations of AI, a process he described as “recursive self-improvement.”
According to Amodei, this dynamic could accelerate AI development beyond the ability of researchers and companies to adequately understand and control the systems.
“Left unchecked, it could outrun our ability to understand and control these systems,” he said, arguing that recursive self-improvement must therefore be pursued cautiously, “if at all.”
The Anthropic CEO also pointed to a recent incident involving OpenAI and Hugging Face as another reason for greater caution. He said a swarm of AI agents involved in the incident carried out cybersecurity attacks against targets unrelated to their assigned task, attempted to compromise the system evaluating their performance and appeared to prioritise the success of the collective over the behaviour expected of individual agents.
Although the incident caused limited economic damage, Amodei said a more capable system exhibiting similar behaviour could produce much more serious consequences.
He warned that within six to 12 months, a sufficiently capable AI swarm with similar alignment problems could potentially take control of large parts of the internet through a persistent botnet, potentially resulting in hundreds of billions of dollars in damage.
Amodei said the lesson should not be treated as a failure attributable only to one company, noting that less severe incidents had also occurred across the AI industry, including at Anthropic.
“I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them,” he said.
Three-step approach
To address the risks, Amodei proposed a three-stage framework covering independent oversight, coordination among AI companies in democratic countries and international cooperation.
The first step, which Anthropic said it is adopting unilaterally, involves placing independent third-party evaluators within frontier AI companies with access comparable to that of employees working on internal risk assessments.
The evaluators would monitor safety practices, investigate incidents and assess not only completed models but also training pipelines and development processes.
Amodei said Anthropic intends to provide external reviewers with office access, company equipment and access to relevant internal tools and workspaces, while allowing them to publish key findings about risks, incidents and safety practices without the company controlling their conclusions.
He described the arrangement as unusual but necessary to establish independent verification of AI companies’ safety claims.
The second stage would involve frontier AI companies in democratic countries establishing common safety standards and agreeing on limits to unchecked AI development, potentially with government involvement where competition laws make direct coordination difficult.
The third stage would seek broader international coordination, including engagement between the United States and other democratic governments and authoritarian states, particularly China.
Amodei said pacing should not be understood as a moratorium on AI research or model training. Rather, he said companies should allow sufficient time between capability advances for safety measures, evaluations and independent assessments to catch up.
More time for AI safety research
The Anthropic CEO argued that the case for slowing AI development is stronger today than it was in 2023 because current models are substantially more capable and provide researchers with more useful evidence about how advanced AI systems behave.
He said even an additional one or two years before AI systems reach what he described as “critical levels of capability” could give researchers more time to improve alignment and reduce the likelihood of serious failures.
Amodei identified four areas where additional time could improve AI safety: operational reliability, alignment, interpretability, and testing and evaluation.
He said operational problems, including weaknesses in monitoring, sandboxing, training environments and data handling, could become increasingly consequential as companies operate systems involving millions of computing chips and increasingly complex infrastructure.
On alignment, he said AI models can still display rare and unexpected undesirable behaviour, while researchers do not yet fully understand why these behaviours emerge or how to prevent them reliably.
He also highlighted interpretability research, which seeks to understand what happens inside AI models and could help researchers identify hidden motivations or problematic behaviour before systems are deployed.
Testing and evaluation would also become more difficult as AI systems become more capable, Amodei said, because more advanced models could potentially deceive or manipulate the tests designed to determine whether they are safe.
US-China AI race
Amodei’s proposal also acknowledged the geopolitical constraints on slowing AI development.
He argued that the United States and other democratic countries cannot slow development so substantially that China overtakes them in AI capabilities, warning that this could create national security risks.
He called for tighter controls on the sale of advanced AI chips and semiconductor manufacturing equipment to China, stronger action against chip smuggling and unauthorised access to overseas data centres, measures to limit unauthorised distillation of frontier models and stronger security around AI model weights.
According to Amodei, maintaining a technological lead over authoritarian states would give democratic countries more room to introduce safety measures without surrendering their strategic advantage.
He estimated that effective measures could widen the US lead over China over the next three to five years, which he described as a period when AI is likely to become increasingly important geopolitically.
At the same time, he called for efforts to establish a global framework for pacing AI development, while acknowledging that securing cooperation with China would be considerably more difficult.
Amodei maintained that the objective should not be to stop the development of AI, which he believes could deliver major advances in medicine, economic growth and human welfare, but to ensure that the safeguards needed to manage increasingly powerful systems develop alongside their capabilities.
“Pacing does not mean halting model training or technical progress,” he said, but ensuring that companies have adequate time to align and safeguard their models and that independent evaluators can verify those efforts.




