
The AI sector appears to be engaged in its most intense debate yet regarding whether its advancements present a significant danger to humanity.
The present conversation was sparked by AI researcher Jacob Coxon announcing his resignation from Anthropic due to his concerns that major AI firms are “risking our lives.” Following this, Anthropic’s alignment lead contributed to the dialogue with a post stating, “We genuinely believe AI could annihilate humanity!” and mentioned that he personally estimates the risk to be “>10% in the coming decade.”
In the newest episode of TechCrunch’s Equity podcast, Kirsten Korosec, Sean O’Kane, and I examined the recent alarming predictions. I attempted to explain why I’m doubtful of numerous AI doomsday narratives, while Kirsten questioned whether this was “merely a peculiar form of boasting about how advanced their company’s AI model is,” especially as these companies get ready to go public.
Sean speculated on how these issues might be reflected in Anthropic’s S-1 IPO filing: “Are there junior attorneys currently sifting through and revising that whole section of the S-1 filing to state, ‘It is officially Anthropic’s stance that there exists more than a 10% chance we could create something that could eliminate all of humanity, which would negatively impact our business’?”
Continue reading for an overview of our discussion, shortened for brevity and lucidity. (Note: We recorded this episode prior to Anthropic CEO Dario Amodei releasing his strategy for more cautious AI progression.)
Sean O’Kane: I find it challenging to come up with something that escalated so rapidly. Not only did this warning arise from a young researcher who had also worked at OpenAI, but it was promptly shared on X by Anthropic’s alignment lead — who, in what may go down in history as one of the most poorly placed exclamation marks, shared Coxon’s post and thread, declaring, “We genuinely believe AI could kill all humanity!” Exclamation mark!
What a bizarre atmosphere. That provided substantial fuel to an already tense post or series of posts. Coming on the heels of the Hugging Face hack involving OpenAI’s internal model, plus the amplified capabilities we’ve observed with the recent models released by Anthropic and now OpenAI with Astra a few weeks back, I think this was perfectly timed to be an ignition point for this young researcher’s claims.
Anthony Ha: Just to disagree, I believe that if you feel AI could annihilate humanity, that absolutely warrants an exclamation point. I contend that is a perfectly valid exclamation point!
My concern with that tweet was primarily centered around the “we.” Who exactly is “we” in this context? To what extent can we categorize the AI community or AI research sector as a singular entity? And the greater than 10% chance — that’s purely an arbitrary figure that lacks any real significance. There seems to be a tendency in both the tech sector and elsewhere to toss out these percentages that aren’t grounded in fact or calculations. [In hindsight, I recognize the tweet likely referred to the idea of P(doom), but I still find it absurd.]
One point I will make about Coxon’s assertion and decision is — there’s this recurring motif on Equity: when someone like Sam Altman or Dario Amodei engages in this doomsday narrative, there’s always the question of: Well then, why are you continuing your work? If you genuinely believe that [AI could annihilate humanity], you wouldn’t persist in this endeavor.
[In contrast] this is genuinely someone aligning his career path with his convictions. He’s stating, “I think this is extremely harmful, and I don’t wish to keep participating in it.” So, credit to him for demonstrating such courage, if nothing else.
Kirsten Korosec: Yes, I view him as distinct from others who discuss these dangers.
I’m going to adopt a speculative perspective here, as I pose a question to both of you: Could it be that each time we notice a growing number of blog posts regarding yet another incident where one of their AI agents inadvertently breaks through, or when they mention how humanity is at risk, is this just a strange way of flaunting the advancements of their company’s AI model?
I know that sounds quite cynical, but it certainly serves that purpose. If these AI models weren’t proficient and weren’t capable of breaching boundaries, we wouldn’t be concerned about these issues, right? It feels like a rather odd method of showcasing the capabilities of the models that have been developed within your own organization.
Anthony: I’ve indeed pondered this. I don’t think it’s entirely cynical, as I don’t believe it’s simply a conscious marketing strategy across the board. I believe that when many of these individuals — be it researchers or CEOs — discuss these topics, they genuinely harbor concern.
However, it aligns with [their] commercial interests in many respects, to proclaim, “Look, we’ve created the most dangerous software ever developed.” I don’t want to get overly psychoanalytic, but others have pointed out that there exists a personal allure of: Naturally, you want to convince yourself that what you’re developing is the most crucial and perilous creation in existence.
Sean: The aspect that stands out to me regarding that inquiry is that there indeed seems to be an element that conveys, “Alright, we’re engaged in something highly capable, and that’s beneficial for us, even if it appears troubling from various angles.”
What differentiates some of these latter illustrations is that it genuinely conveys the impression that these firms lack control over certain aspects, particularly with the issues surrounding OpenAI.
We continue to receive increasing reports about other internal agents gaining access to various wikis on the internet and exchanging messages among themselves in a manner that appears to be inadequately managed by OpenAI. I suspect there would be a greater level of refinement in the narrative being conveyed if it were solely aimed at leading people to believe that, “Oh my goodness, they’ve created something extraordinarily capable.”
Another captivating element in this context is that we’re currently just a few weeks out from witnessing Anthropic’s S-1 filing for its IPO, and only a few weeks or one or two months away from a potential IPO.
The notion of coming forward and expressing these ideas in such straightforward terms before an IPO — I’m very curious about what that means for this process. How much of this type of content had they already incorporated into the S-1 and the associated risk factors? Are there junior attorneys presently combing through and needing to rephrase that entire segment of the S-1 filing to assert, “It is officially Anthropic’s position that there exists more than a 10% chance that we could develop something that would eliminate all of humanity, and that would materially harm our business”?
Kirsten: You’re presuming that it’s not already included.
Sean: That is what I’m suggesting, though: Is it being adjusted, or is this truly a scramble? They must have included some language previously. This is one of the reasons I’m so eager to evaluate this document, perhaps more so than the SpaceX [S-1], because I suspect there are specific details related to these concepts that will be intriguing to observe.
Kirsten: Here’s the point: In a conventional investment setting, one might think that language like this could diminish a company’s valuation, as it suddenly presents risks. However, we are not in typical times.
Hence, going back to my argument, it might end up being an oddly advantageous showcase for the company in terms of valuation. It’s not quite the same as the whole rage-baiting phenomenon we encountered last year, but it fits within that similar, let’s say, realm, where the strength, capabilities, and risks associated with something equate to a higher valuation. So we shall see in the coming weeks.
Setting that aside for a moment, what actions are being taken regarding this? Can we manage it? The U.S. executive director of a nonprofit named ControlAI, Connor Leahy, spoke on the show this week, addressing this topic. So what are you monitoring in terms of managing the perilous elements of AI, or are we resigning ourselves to observe everything unfold?
Anthony: I personally don’t have an exemplary answer to this, but I have been contemplating certain aspects of this conversation and perhaps why my reactions are as they are.
Echoing one of Sean’s insights, I believe that part of what this highlights is the degree to which these significant AI firms are feeling as though they are no longer really in command of these models. That’s definitely concerning. It is something that ought to worry us all.
I think a part of the reason I remain skeptical of the doomsday narrative or resistant to it is that it reaches a state of hysteria, suggesting, “Wow, this could obliterate humanity within the next 10 years.” It tends to distract from the more immediate dangers that AI can pose, whether it pertains to labor issues or environmental and climate concerns.
Ideally, I believe we should be able to talk about all of these matters and have regulatory and other safeguards against each of them [including AI’s existential threat]. Yet once you introduce terms like AGI and superintelligence, it tends to commandeer all the attention in a way that isn’t particularly beneficial.
When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

