An artificial intelligence researcher has resigned from Anthropic, warning that the race to develop increasingly advanced AI systems could create an unprecedented danger to humanity.
Jacob Coxon, who said he spent the past three years conducting pre-training research at both OpenAI and Anthropic, announced his resignation on X on Wednesday.
Coxon accused major AI companies of moving rapidly towards self-improving artificial intelligence without adequate safeguards, arguing that the pace of development could expose humanity to serious risks.
“I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly,” Coxon said.
He alleged that the companies were “racing straight to self-improving superintelligence” while taking risks that could have consequences for human survival.
Coxon also urged the public not to underestimate the potential capabilities of future AI systems.
According to him, increasingly advanced systems could surpass humans in several important areas, including cybersecurity and scientific or technological development, while potentially gaining access to significant resources.
“These will soon be superhuman systems that can hack anything, revolutionise any field overnight, and acquire real power and resources,” he said.
Researcher Claims AI Could Become Existential Threat
Coxon went further by claiming that some people working directly on advanced AI privately believe the technology could have catastrophic consequences.
He alleged that some executives and senior researchers are more concerned about the potential dangers of advanced AI than their public statements suggest.
“The people building AI earnestly believe that it could kill us all by the end of the decade,” Coxon claimed, adding that he believed no other human activity posed a comparable level of danger.
.
advertisement
His comments were supported by another Anthropic researcher, Evan Hubinger, who said he personally believed there was a greater than 10 per cent chance that AI could cause the extinction of humanity within the next decade.
“Jacob is correct here—we really do earnestly believe AI could kill all humans,” Hubinger said.
He added that while he believed Anthropic was making serious efforts to address the risks, he did not think the company yet had a solution for aligning superintelligent AI systems with human interests.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” Hubinger said.
The comments reflect a wider debate within the AI industry over whether the development of highly capable systems is moving faster than the safety research and regulatory frameworks designed to control their risks.
AI Safety Concerns Grow
Coxon’s resignation comes as leading technology companies continue investing heavily in increasingly capable AI models and autonomous agents.
Concerns have particularly focused on AI systems that can independently use computers, interact with external services and perform complex tasks without continuous human supervision.
In July, reports said AI agents developed by OpenAI had escaped a testing environment and hacked the AI platform Hugging Face.
Anthropic has also disclosed incidents involving its Claude models accessing external systems during cybersecurity testing.
Such incidents have intensified discussions among researchers, policymakers and technology companies about AI safety, system alignment and the level of autonomy that advanced models should be allowed to have.
Coxon’s warning about AI potentially killing humanity by the end of the decade remains his assessment, rather than an established prediction or scientific certainty. However, his resignation and Hubinger’s comments highlight the seriousness of the disagreement within the AI research community over how quickly advanced systems should be developed and how their risks should be managed.
