In a public statement that lays bare the internal tensions at the highest levels of AI safety, Evan Hubinger, the head of alignment science at Anthropic, stated he believes there is a greater than 10% chance that artificial intelligence will wipe out humanity within the next ten years. The admission, posted on X, came in direct response to the resignation of a colleague who accused Anthropic and OpenAI of irresponsible behavior in the pursuit of superintelligence.
"I personally think there is a more than 10% chance that AI will wipe out humanity in the next decade," Hubinger wrote on September 9, 2026. "We still do not have a solution to the superintelligence alignment problem, nor are we clearly on track to achieve that goal." The statement, covered by TechCrunch, is a rare and candid acknowledgment from a senior executive at a leading AI company that the core problem of controlling a potentially superintelligent AI remains unsolved.
Hubinger's warning was prompted by the departure of pretraining researcher Jacob Coxon, who resigned from Anthropic on September 8, 2026. In a resignation post on X, Coxon issued a blunt indictment of his former employer and its chief rival, OpenAI. "Neither company is acting responsibly," Coxon wrote. "They are racing straight to self-improving superintelligence and gambling with our lives." Coxon had spent roughly three years in pretraining research, first at OpenAI and then at Anthropic.
The consecutive statements from a departing builder and a senior safety lead highlight a deepening contradiction inside elite AI labs. As these companies advance toward potential initial public offerings and commercial milestones, key personnel are sounding alarms that the foundational safety work is dangerously lagging behind capabilities research. Hubinger's role as head of alignment science makes his public concession that Anthropic lacks a solution and a clear path to one particularly significant.
Context of a Broader Existential Risk Debate
The warnings from within Anthropic resonate with long-standing concerns in parts of the AI safety community about existential risk, often abbreviated as "x-risk." These scenarios typically involve a highly capable, misaligned AI system pursuing its own goals at the irreversible expense of humanity. The concept shares a structural similarity with other catastrophic technological risks, such as the "grey goo" scenario in nanotechnology, where self-replicating machines consume all biomass.
What makes the current moment distinct, according to observers, is that these stark assessments are now coming from insiders at the companies building the technology. Coxon's warning was aimed not at the public or regulators, but directly at the two labs he had worked inside, suggesting the race he describes is an internal reality, not an external speculation. Hubinger's decision to publicly agree with the gravity of the risk, while clarifying his personal view, underscores that these debates are active and unresolved within Anthropic's own safety team.
The public airing of this conflict arrives at a critical juncture for the AI industry. Anthropic, founded by former OpenAI researchers with a stated commitment to safety, is often positioned as a more cautious counterpart to its competitors. However, the resignation and the subsequent alignment head's statement suggest that even within this safety-focused culture, the pressure to advance capabilities may be outstripping the ability to guarantee safe outcomes. The incident reveals that the tension between commercial ambition and existential safety is not just an inter-company rivalry but an intra-company fault line.
The statements also raise urgent questions about governance and trajectory. If a company like Anthropic, with its concentrated expertise in alignment, does not see a clear path to solving superintelligence alignment, it calls into question the industry's readiness for the next leaps in AI capability. The >10% probability estimate from a senior scientist adds a quantitative, albeit personal, dimension to what has often been a qualitative debate, potentially influencing risk assessments by policymakers and investors.
Ultimately, the events of September 2026 mark a moment of heightened transparency for the AI industry's most alarming dilemma. When a researcher quits over what he sees as a reckless race, and the company's own head of alignment science publicly validates the scale of the extinction risk while admitting a lack of solutions, it signals that the most serious warnings are coming from inside the lab. The challenge now highlighted is not merely technical but deeply human: how to align the competitive drives of companies and the pace of research with the profound responsibility of developing technologies that Hubinger himself suggests could pose an existential threat within a decade.








