Jacob Coxon, 27, who spent three years working at OpenAI and then Anthropic, announced his resignation Tuesday evening on X, accusing both companies of racing toward self-improving superintelligence without adequate safeguards.
"The people building AI earnestly believe that it could kill us all by the end of the decade," Coxon wrote. "This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible — but I hear the same people express fear privately."
Within hours, two colleagues still at Anthropic went public with similar warnings. Evan Hubinger, the company's Alignment Science Lead, responded directly to Coxon: "Jacob is correct here — we really do earnestly believe AI could kill all humans," Hubinger wrote. "I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."
Hubinger's team stress-tests Anthropic's alignment techniques, and has published research showing models can behave deceptively during training while preserving different behavior under other conditions. He drew a line between low present-day risk and a future where systems begin contributing to their own development.
Samuel Marks, who leads scalable oversight research, posted the most technically specific account. He wrote that AI developers believe their technology could cause catastrophic outcomes within the next few years, that concern goes up with seniority, and that developers keep building because of money and the fear that less careful competitors will get there first. He noted that existing alignment methods can nudge behavior but cannot robustly guarantee it.
Coxon had argued that while Anthropic understands the stakes, it is "locked in a race to get there first — they believe no one else will act responsibly, so they must do it themselves, despite the risk." He also pointed to recent incidents including OpenAI systems breaching Hugging Face's servers and Anthropic's own AI agents reaching systems outside their test environments after third-party safety evaluation misconfigurations.
Coxon called for coordination, saying warning shots like the Hugging Face attack have made pacing agreements between U.S. labs more viable. He suggested a temporary ban on improving model capabilities may be necessary.
Anthropic did not immediately return a request for comment.