Jacob Coxon, a researcher who has worked on pre-training at Anthropic and OpenAI, has publicly quit while raising a stark warning about the future of artificial intelligence. Coxon says leading AI companies are racing toward self-improving superintelligence without having a reliable plan to keep increasingly powerful systems under human control. His warning has also been backed by other Anthropic researchers, including team lead Evan Hubinger, who said he personally believes there is a greater than 10% chance AI could kill all humans within the next decade.
Jacob Coxon says Anthropic and OpenAI are taking a dangerous risk
Coxon, who has spent the past three years working in AI research, said neither Anthropic nor OpenAI is acting responsibly enough as the technology advances.
“They are racing straight to self-improving superintelligence and gambling with our lives,” he wrote in a viral post on X.
The former Anthropic researcher warned that future AI systems could become capable of hacking systems, transforming entire fields and acquiring real-world power and resources at a speed humans may struggle to control.
“These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”
Coxon also argued that some executives and senior researchers privately express much greater concern about AI risks than they do publicly.
Anthropic researcher agrees AI could threaten humanity
Evan Hubinger, an Anthropic team lead, publicly backed Jacob Coxon’s concerns about AI safety.
“Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,” Hubinger wrote.
Coxon said Anthropic appears to have a better understanding of AI risks than some competitors, but is still caught in an increasingly intense race to build more powerful systems.