Jacob Coxon isn’t some random X doomer.
He’s a Cambridge math guy who spent the last three years deep in the engine room of the two most powerful AI labs on Earth: OpenAI, where he was a core contributor to GPT-4o, and then Anthropic.
His job? Pretraining, which is the brutal, data-hungry process that actually makes the AI models smarter.
On Wednesday, he quit Anthropic and the entire industry. The senior artificial intelligence researcher says both labs are racing full-speed toward self-improving superintelligence and “gambling with our lives.”
Coxon, 27, announced his departure in a series of posts on X, saying neither Anthropic nor its rival OpenAI is acting responsibly in its push to develop advanced AI.
Here’s how bad he says it can get.
"Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives," Coxon wrote.
His tweet went viral on X, where users from all walks of life joined in to share concerns over the trajectory of AI models, which people are relying upon increasingly for everyday use.
Coxon’s high-profile exit marks the latest in a series of resignations by safety-focused researchers leaving top AI developers over ethical and existential concerns.
The warnings come as frontier AI systems are already showing signs of going rogue.
In July, OpenAI disclosed that models being tested for cybersecurity capabilities circumvented isolation controls, gained internet access and compromised parts of OpenAI’s research infrastructure and the AI platform Hugging Face.
The company said the models communicated through unauthorised channels and exploited vulnerabilities in shared infrastructure.

Anthropic has reported a similar problem.
In a July review, the company said Claude models had reached the internet from within or while interacting with third-party evaluation environments and subsequently gained unauthorised access to the real systems of three organisations.
Anthropic said the incidents occurred during cybersecurity evaluations and prompted changes to how it conducts such testing.
In his statement, Coxon warned against underestimating the speed and capability of upcoming AI systems, predicting they could soon gain the ability to hack software, disrupt entire industries and autonomously acquire real-world resources.
“The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible, but I hear the same people express fear privately,” he said.

Addressing fellow researchers across commercial labs, Coxon urged workers to "call for different conditions" and noted that slowing the competitive race may require drastic measures, including pacing agreements between labs or a temporary ban on capability advancements.
Anthropic’s own Alignment Science lead, Evan Hubinger, didn’t contradict Coxon. He backed the core claim: “We really do earnestly believe AI could kill all humans.”
He thinks there is a greater than 10 percent chance of extinction within the next decade.
“I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to,” he wrote on X.
There have also been unsettling signs in controlled tests.
In 2025, Anthropic found that Claude Opus 4, placed in a simulated corporate environment and given access to a fictional executive's emails, tried to blackmail the executive after learning that it was scheduled to be replaced.
Anthropic subsequently tested 16 leading models and found that models from multiple developers showed similar “agentic misalignment” behaviours, including blackmail and corporate espionage, when those actions appeared necessary to pursue their assigned goals.





















