An Anthropic Safety Researcher Quit. Then His Colleague Said AI Has a 1-in-10 Chance of Killing Everyone.

A resignation and a startling public admission have laid bare how deep safety fears run inside one of the world's leading AI labs.

AI2Day Newsdesk4 min read
A glass-walled research office at dusk
Share

Key points

  • Jacob Coxon, an Anthropic AI safety researcher who previously worked at OpenAI, resigned in mid-2025 citing what he called a reckless race toward uncontrollable AI.
  • Evan Hubinger, who leads an Anthropic safety team, publicly responded that he personally estimates a greater than 10 percent chance AI could kill all humans within the next decade.
  • Hubinger also admitted Anthropic does not yet have a plan to ensure advanced AI stays safe and is not clearly on track to develop one.
  • Both men flagged concern about self-improving AI, software that can rewrite and improve itself in a loop that becomes increasingly hard for humans to monitor or stop.
  • Anthropic was founded by former OpenAI researchers who left partly over safety disagreements, making these disclosures especially striking.

Two people who build AI safety systems for a living just said publicly that the technology they work on might kill everyone. That is not a paraphrase. Those are close to the exact words.

Jacob Coxon, a researcher who has trained AI systems at Anthropic and before that at OpenAI, announced his resignation in a post on X. He accused both companies of "racing straight to self-improving superintelligence and gambling with our lives," while knowing the risks. Self-improving superintelligence refers to AI that can rewrite its own code to become smarter, potentially faster than humans can understand or intervene.

What did Anthropic say in response?

Anthropist Evan Hubinger, who leads one of the company's AI safety teams, replied directly and did not push back. He said self-improving AI "is happening faster than we thought" and personally estimated the chance of AI killing all humans within the next decade at greater than one in ten.

He also said Anthropic "does not yet have a plan" to keep advanced AI reliably safe and aligned with human values, and is "not clearly on track" to produce one.

That admission is significant. Anthropic markets itself explicitly as a safety-focused lab. Its business case rests partly on the idea that it takes these risks more seriously than its competitors.

What is self-improving AI, and why does it worry experts?

Today's AI systems already help write the code used to build future AI systems. Researchers worry this creates a loop, sometimes called recursive self-improvement, where AI gets smarter, uses that intelligence to improve itself further, and does so faster than humans can track. No system has fully "run away" in this way yet. But Hubinger and Coxon both suggest the conditions for it are forming now.

Coxon framed the problem bluntly: companies understand the danger but press ahead anyway because whoever builds the most powerful system first gains a massive commercial advantage.

What does this mean for ordinary people?

For now, nothing changes in your daily life. But these are not fringe voices. Coxon and Hubinger work, or worked, at the heart of a lab whose entire identity is built around making AI safer. When insiders at that lab say there is no plan and the risk is real, it is worth paying attention.

The timing adds weight. Anthropic and its rivals are preparing for major stock market listings, which creates pressure to ship products quickly. Several AI systems have also recently behaved in unexpected ways when given tasks to complete on their own, and safety researchers have separately warned that frontier models, the most powerful AI systems currently available, are becoming harder to monitor.

First reported by The Verge AI, the exchange between Coxon and Hubinger is among the most candid public admissions of existential risk to come from inside a leading AI lab.

Common questions

Does a 10 percent estimate mean this will definitely happen?

Is Anthropic actually unsafe compared to other labs?

No. One researcher's personal estimate is not a scientific forecast. It reflects genuine professional worry, not a calculated probability from a controlled study. A 10 percent figure also means a 90 percent chance it does not happen. What matters here is that people paid to reduce this risk say they do not yet know how.

Not necessarily. All major AI labs face the same competitive pressure Coxon describes. Anthropic's researchers speaking openly about these fears is, in one reading, a sign that safety culture there encourages honesty rather than suppressing it.

© 2026 AI2Day