A heated debate has intensified within the artificial intelligence industry over the possibility that the technology could become an existential threat to humanity, following researcher Jacob Coxon’s resignation from Anthropic because he said he believed that the largest AI companies were “gambling with our lives.” After that, Anthropic’s head of alignment shared a post saying that the company seriously believes AI could kill all humans, and estimated the probability of that happening at more than 10% over the next decade.
These developments were the focus of an episode of TechCrunch’s Equity podcast, featuring Kirsten Korosec, Sean O’Kane, and Anthony Ha. The article, an edited excerpt from the discussion, does not provide an independent scientific estimate of the probability of catastrophe. Instead, it examines the motivations behind this language and its effects on companies, investors, and the public.
A Big Number Without a Disclosed Methodology
Anthony Ha questioned the use of the figure “more than 10%,” considering that, based on what emerged in the discussion, it was not supported by an explained calculation or methodology. The participants linked it to the concept of P(doom), a term used to describe the probability of an AI-related catastrophe. But stating a specific percentage does not automatically make it a verifiable estimate, especially when the assumptions or data on which it was based are not provided.
By contrast, Ha saw Coxon’s resignation as putting him in a different position from executives and researchers who warn about AI risks while continuing to work on developing it. The decision to leave his career path, as the participants discussed, represents an action consistent with his belief that the risks are severe, even if it does not prove that his estimate of the threat is correct.
A Genuine Warning or a Display of Capabilities?
Kirsten Korosec raised a more skeptical possibility: repeated warnings about AI agents escaping control or posing a danger to humanity may serve an indirect marketing function. Suggesting that a model is so advanced that it inspires fear can also be read as an announcement of its capabilities, and may serve the interests of the companies developing these models, even if the stated concerns are sincere.
Ha did not reduce these statements to a deliberate marketing campaign. In his reading, researchers and CEOs may have genuine concerns, but those concerns also intersect with a commercial interest: presenting the technology as among the most important, influential, and dangerous forms of software in history. It is therefore difficult to separate the warning from the personal and institutional incentives surrounding it.
The Control Problem Goes Beyond Catastrophic Rhetoric
Sean O’Kane noted that some recent events may give the impression that companies do not have complete control over their systems. The discussion cited what was described as a breach connected to an internal OpenAI model, cases in which internal agents accessed web-based wikis and left messages for one another, and the leap that the participants believed had appeared in the capabilities of Anthropic and OpenAI models, including the Astra system, which was launched weeks before the episode was recorded.
These events, as presented in the episode, do not prove that AI is capable of eliminating humanity, but they raise a more immediate practical question: Can companies reliably test, monitor, and contain the behavior of their agents? This is an important point for technical readers; the gap between a model’s capabilities and an organization’s ability to manage its behavior may be more observable than distant predictions about artificial general intelligence or superintelligence.
What Could This Mean for Anthropic’s IPO?
O’Kane discussed the possibility that these statements could affect the registration statement Anthropic is expected to file for an initial public offering, known as an S-1, and its risk-factors section. The participants wondered whether the company had already included language concerning these risks or had been forced to reword it after the recent statements. The source did not provide an answer, and the filing or offering date had not been settled in the article.
Existential-risk language may appear negative in a traditional investment environment, but Korosec raised the opposite possibility: some investors may interpret high capabilities, and even elements of danger, as evidence of the company’s value rather than as a burden on it. This leaves the valuation question open until the company’s actual documents emerge, particularly its risk disclosures and operating assumptions.
Editorial Reading
The importance of the discussion does not lie in proving the “more than 10%” figure, since the source provides no basis for verifying it. Rather, it lies in revealing the tension among three simultaneous narratives: companies saying that the risk is severe, companies continuing to develop models and market their capabilities, and investors who may read the danger itself as a signal of value. Ha warns that focusing on extinction scenarios may displace more near-term harms, such as AI’s effects on jobs, the environment, and the climate. At the same time, this does not eliminate the need to examine existential risks and regulatory safeguards; instead, it requires discussing them alongside current harms and with clearer evidence than unexplained percentages.
It is worth noting that the episode was recorded before Dario Amodei, Anthropic’s CEO, published a plan to develop AI more cautiously; therefore, the source does not cover the content of that plan or its effect on the discussion.