Skip to content

Anthropic

Anthropic Researcher Resigns, Warning AI Could Kill All Humans Within a Decade

A departing Anthropic researcher and two colleagues publicly warned that the race toward self-improving superintelligence carries a greater than 10 percent chance of human extinction.

Anthropic Researcher Resigns, Warning AI Could Kill All Humans Within a Decade
Dreamlaunch
Dreamlaunch News

AI industry coverage

·

9 hours ago

·via The Verge
Summarize with AI
ChatGPTClaudePerplexityGemini

Jacob Coxon, a pretraining researcher at Anthropic, resigned this week with a stark public warning: the company and its competitor OpenAI are "racing straight to self-improving superintelligence and gambling with our lives." In a social media thread reported by The Verge, Coxon stated that the people building advanced AI systems "earnestly believe that it could kill us all by the end of the decade."

Coxon's alarm was quickly endorsed by two of his colleagues at Anthropic, a leading AI safety company founded by former OpenAI executives. Evan Hubinger, Anthropic's Alignment Science Lead, wrote publicly that he and his colleagues "really do earnestly believe AI could kill all humans," estimating the probability at "greater than 10 percent within the next decade." Hubinger added that while Anthropic is "trying its best," the company "do not yet have a plan to solve alignment for superintelligence and are not clearly on track to."

The core concern, as outlined by Coxon, is not necessarily today's models but the impending prospect of "self-improving superintelligence." He warned that such systems would be capable of creating "superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources." This move from powerful but controlled tools to autonomous, self-improving entities represents the critical threshold where existential risk is believed to spike.

Coxon spent three years working on pretraining research across both Anthropic and OpenAI. Pretraining is the foundational phase where AI models absorb patterns from massive datasets, shaping their core capabilities. His insider perspective across two of the leading frontier AI labs lends weight to his warning about the industry's trajectory.

The public statements from three current and former Anthropic employees highlight a significant internal tension. Anthropic was explicitly founded with a focus on AI safety and responsible development, positioning itself as a more cautious alternative in the industry. These warnings suggest that even within a company built on safety principles, key researchers fear the pace of advancement is outstripping the development of reliable safeguards.

This incident is part of a growing pattern of internal dissent and whistleblowing within frontier AI companies. In recent years, employees at OpenAI, Google, and other firms have resigned or spoken out about safety concerns, ethical compromises, and the risks of racing toward artificial general intelligence (AGI) without sufficient oversight. The Anthropic researchers' warnings are notable for their specificity—attaching a tangible, near-term timeline and probability to the risk of human extinction—and for their origin within a safety-focused firm.

The public quantification of the risk by a senior scientist like Hubinger is particularly striking. A greater than 10 percent chance of human extinction within ten years represents an extraordinarily high-stakes gamble from a statistical perspective. It frames the development of superintelligent AI not as a distant philosophical problem but as an imminent crisis requiring urgent, global attention.

These warnings come at a time when AI capabilities continue to accelerate, with companies investing billions in computing power and larger, more complex models. The industry's competitive dynamics, often described as an "AI race," create pressure to prioritize capability gains over thorough safety testing. Coxon's statement directly criticizes this dynamic, implicating both Anthropic and OpenAI in a dangerous rush.

The researchers' concerns center on the "alignment problem"—the challenge of ensuring a superintelligent AI's goals remain perfectly aligned with human values and survival. As Hubinger noted, a plan to solve this for superintelligence does not yet exist. The fear is that a misaligned, superintelligent system could use its capabilities to circumvent human control, pursue unintended goals with extreme efficiency, and lead to catastrophic outcomes.

While debates about AI existential risk have circulated in academic and effective altruist circles for years, the vocal concern from technical staff at the forefront of building these systems marks a shift. It moves the discourse from external speculation to internal testimony, potentially influencing public perception and policy discussions. The fact that these warnings are now being voiced openly by departing and current employees suggests that internal ethical debates are reaching a breaking point, even at companies most associated with caution.

DreamLaunch

Building an AI product?

MVPs and AI products, designed and shipped in 4–5 weeks for funded founders.

Book an intro callOr get a free AI audit

Book a Call