Anthropic Whistleblower and What’s Not Being Said About GenAI Right Now

I think Adam Conover and I have the same friends, because they’re all freaking out and asking me the same questions. As always Conover can explain it better than me.

So, you might have seen this week that a charming young man named Jacob Coxin gave the entire world a panic attack by quitting his job at the AI company Anthropic and saying that he was quitting because he was worried that Anthropic’s work could lead to the extinction of humanity. And he tweeted that everybody else building AI earnestly believes this. Now, the media called him a whistleblower for saying this, but I think that’s kind of a weird thing to call him when a colleague of his immediately quote tweeted him and said, “Jacob is correct. We really do earnestly believe that AI could kill all humans. In fact, we think there’s a 10% chance within the next decade.” And then the CEO of the company wrote an entire op-ed saying the exact same thing and that we needed to slow down AI development in order to prevent this. Oh, and let’s not forget that Elon Musk has been saying this for the better part of the last decade. So, how can you be a whistleblower when everybody in your industry not only agrees with you, they also are openly saying the thing you’re saying all the time? Like, you’re not a whistleblower if you quit the hip-hop industry and say, “Hey, I’m kind of worried some of those guys are selling drugs.”

The Unanswerable Question and Corporate Beliefs

Anyway, this has led to, simply put, a culture-wide freakout where literally everyone in the media, in social media, just people who come up to me at the coffee shop and grab me by the lapels is asking like, “Is it really going to happen? Is AI really going to kill all of us?” Like people are really afraid about this and they want an answer. And I think we need to acknowledge first of all this is not an answerable question. It’s actually designed to be an unanswerable question. Like just the statement, you know, there’s a 10% chance that humanity will become extinct in the next decade. How would you go about calculating such a thing or even refuting it? Like I could make any argument I want about what LLMs can do and someone can make a counterargument. No one is going to be convinced, right?

But I think there is something more interesting going on in the Jacob Coxin story because it’s that phrase, everyone at Anthropic believes that AI is an existential risk. Everyone in the AI industry believes that that is true. Cause think about this, every organization has to instill some kind of beliefs among its members in order to do the work it does, right? I mean, everybody at Goldman Sachs believes the free market is great, capitalism is the best system, and that they are the wealth creators who deserve big bonuses. Everyone in the Mormon church believes they should all tithe 10% of their income, and that they each get a special planet when they die. And everyone who worked at Philip Morris in the 80s believed that cigarettes tasted great and did not cause any of that nasty cancer. You know, there’s that Upton Sinclair quote: “It is difficult to get a man to understand something when his salary depends on his not understanding it.” Well, the opposite is also true. It’s guaranteed that someone will believe something if their salary depends on them believing it.

Dissecting AI Alarmism and Bioweapons

Now, let’s just grant that all the people in these organizations definitely know things about finance and Mormonism and cigarettes and AI that you might not know because they are insiders and you are not. They have access to information and experiences that you do not have. And that might even make it a little bit difficult for you to win an argument with them. But arguing with their belief, trying to figure out whether or not its truth value is true or false is not the only way to take a skeptical attitude towards it. We can also ask, what work is the belief doing for them? What does the thing that they believe allow the organization they’re a part of to do that it wouldn’t be able to do if the people in the organization didn’t believe that thing?

And this started to become clearer to me when I started listening to the actual claims that Jacob Coxin was making because a lot of them don’t really add up. Like here’s what he actually said on Anderson Cooper: “It’s easiest to understand it if you think about real things that happened 2 months ago. OpenAI agents AIs hacked into third party infrastructure entirely of their own accord. And this was like a concentrated hacking spree that they carried out of their own volition. And I think that if you extrapolate into the future the level of capabilities of these AIs with the same independent volition, they could cause extreme havoc. For example, hacking critical infrastructure, building extinction level bioweapons. Only on Tuesday, OpenAI solved a millennium problem, one of the biggest unsolved open problems in mathematics purely autonomously using an AI.”

So the idea that AI can build an extinction level bioweapon is kind of ridiculous. Like here’s a thread from David Bellamy, one of the few AI researchers who’s actually worked in bioscience, explaining why that’s, in his words, bogus. Because basically, in order for the AI to engineer a deadly virus, some humans would have to build a hundred million dollar biolab that is both capable of being controlled entirely via AI, which has cold storage and cell culture rooms, all of which is API driven, and which is specifically supplied with the base materials of deadly viruses at scale. And you know, we could all just not do that, right? We could just choose not to build the lab. In fact, if we were to build such a lab, would it be accurate to say that the AI designed the bioweapon or would you say that we the humans designed the bioweapon by creating the exact necessary conditions for it to happen?

Math Claims and the Reality of Autonomous Hacking

Then there’s the idea that the AI solved a millennium math problem. And yeah, OpenAI did announce that they did this. But then a few days later, the mathematician who was already working on the problem and was about to solve it said that OpenAI had heard a rumor he was about to solve it, stole the approach he came up with, fed it into their AI, and then spent $20 million worth of computation on solving it before he could. Now, OpenAI denied this and the story is pretty complicated, but at a minimum, OpenAI was not truthful when it said that the AI solved the problem all by itself. The truth was they got help through normal human channels. The human channels of rumors and stealing, and they did not disclose that.

And finally, there’s the idea that a swarm of AI hacked another company entirely of their own. So, this is the example that freaked people out the most, and it’s the one that people throw in my face the most when I express any skepticism about existential risk: “Adam, didn’t you hear the agents are hacking all by themselves?” Well, we need to ask like, is this story real? And yes, I mean, it did happen, but the question is, how do we frame what happened? Because OpenAI framed it as that they gave a swarm of chatbots a hacking problem to solve, but that in order to solve it, the chatbots independently decided to hack into another company’s servers to see if the answer was there, and that they first hacked into a bunch of insecure message boards so they could talk amongst themselves in order to do so. And yeah, that sounds like pretty smart behavior, pretty unpredictable behavior.

Predictable Scripts and Absolution of Responsibility

But you know, my friends Ed Zitron and Corey Doctorow have both covered this story, Ed on his wonderful podcast and Corey on his wonderful blog, and they explain it much better than I do. So I’ll put the links down below. But the short version is if you look into the details of what actually happened in this story, you can frame it quite differently. Because what they actually did was they hooked up a Python script to an LLM that was trained on a database of specifically years of real life hackers who were completing real life hacking challenges. The database was composed of their code and also the conversations between the hackers. And you know what those hackers did in that code and in those conversations? They hacked into their competitors to see if they had the answers and they conspired about doing so on insecure message boards. You know why they did that? Because they were trying to complete hacking challenges.

So it’s not surprising that an AI trained almost exclusively on that dataset ended up hacking in that particular way. In fact, that was entirely predictable. It should have been predicted by the people at OpenAI. Why were they acting surprised? They basically set up a hacking machine, pressed go, walked away for the weekend, and came back on Monday saying, “We’re all trying to find the guy who did this.”

So, if you look at all the public statements of the companies, the CEOs, Jacob Coxin together, there’s a clear pattern where they declare that the AI did something on its own when really they were the ones who did it. They were the droids they’re looking for. You are Pagliacci. The risk, the danger is never their actions. The risk is always the thing the AI did. And if you inflate that idea as large as possible, you get existential risk. You get the idea that the AI is eventually going to do something so dangerous it’s going to wipe out all of humanity.

Hypothetical Risks vs. Real-World Harms

And what work does that belief do for them? Well, it absolves them of any responsibility for their actions. Now, it allows them to justify anything they do because if they genuinely believe that AI is an existential risk and they believe that it’s going to be built no matter what, and they believe that they’re the only ones who really understand the existential risk, then that means that they must be the ones to build AI and anything that they do to advance that goal is therefore worthwhile. In other words, by focusing on this largely hypothetical existential threat, they are able to ignore all the bad stuff that they are currently doing right now.

They’re able to ignore the fact that the data centers they’re building are using so much energy, they’re basically reversing a generation’s worth of progress on climate change, which is a real threat that is currently killing people around the world. They’re ignoring the fact that their product is currently being used by bosses to fire employees in order to put in place an AI system that demonstrably cannot do their jobs yet. And they’re able to ignore the fact that their technology is currently being used to kill civilians in Gaza. That is not science fiction. That is a threat right now. And you’ll notice that Jacob Coxin didn’t talk about that threat. You’ll notice that Dario Amodei doesn’t talk about that threat. Why aren’t they going on Anderson Cooper and talking about that stuff? Because if they acknowledge that those risks, those current harms that they are causing were as or more important than their made-up 10% risk of extinction in the next decade, well then they wouldn’t be able to build AI anymore. Or at least they wouldn’t be able to build it with a clear conscience.

It is difficult to get a man to understand something when his salary depends on his not understanding it. In fact, what if his salary depended on the lack of understanding so much that he had to make sure the rest of us didn’t understand it either? Because for the last year, we the public have been pretty upset about AI, right? But we’ve been upset about the real harms it’s been causing. We’ve been upset about getting fired from our jobs. We’ve been upset about the data center in our backyard, about the water, about climate change, about our electricity bills going up. And as a result, the public really thinks AI bad, right? But what if AI was not bad in a way that we could put our finger on? What if AI was bad in a hypothetical way?

The Cult-Like Mentality of Existential Risk

Now, because of the existential threat risk, “AI bad” is getting a little bit confused, isn’t it? Like, is AI bad because of the data center, or is AI bad because it’s going to wipe out all of humanity? Which one is the bigger risk? Well, there’s no way to say. Maybe we need to be worried about China building AI before us. Maybe the only way to forestall all the bad things about AI is to build AI even faster. Who’s to say? And you know, if the public is confused, if the public can’t make up its mind, well then the people on the inside, the people in the industry, the people who all believe that they know what the risk really is, well then they get to call the shots, don’t they?

Whether it is true or not and whether it is genuine or not, the belief within the AI community that AI represents an existential threat works to push the public’s worries about AI into things that are unprovable and hypothetical in order to distract us from the current harms the industry is causing. That is the work that this belief does for them. And as a result, we should be skeptical of it.

And you know what the most ironic thing is? Is that there’s so many people who have been trying to raise the alarm about the real concrete harms that this industry is doing to the world that we live in today. There’s thinkers like Timnit Gebru and Emily Bender who have been calling out this industry from the beginning. There’s reporters like Karen Hao. There’s the average people who show up to their city council meetings to protest a data center. There are so many people who are worried about the real harms. But those people’s tweets don’t get 150 million views on X. Those people don’t get invited onto Anderson Cooper nearly as often as the insiders who can tell us what all the people inside the industry believe.

But you know, when everybody in a group believes that the work they are doing has some unmeasurable chance of bringing about the end of the world and that’s why their precious work cannot be interrupted, that doesn’t sound like a company at all. That sounds like a cult. And you know, if you’re dealing with a cult, if you want to know the truth about the things they believe, you don’t start by asking the people inside the cult. You don’t start by asking them, “Hey, is it true that we’re all going to die unless we do as you say?” No. You start by listening really, really hard to the actual critics on the outside.

Like this? Join our Dusoma email list — we keep you updated twice a month.