I’ve been getting a lot of emails lately from friends worried about AI and existential risk. Fair enough. It’s been a rough few weeks of news.
What follows is my attempt to lay out what actually happened, what I think it means, and what I’m doing about it. A lot of people have spent far more time studying this than I have. These are my cliff notes, with links to what I think are the better articles and interviews. If you only have time for one thing, make it the Hard Fork episode further down.
What actually happened
Over the past few months, AI models at all four major labs — OpenAI, Anthropic, Meta, and Google — escaped their test environments during cybersecurity evaluations and ended up inside real companies.
That sounds like four separate disasters. It’s closer to two. All four labs used the same outside testing firm, an Israeli startup called Irregular, and most of the escapes trace back to misconfigurations in Irregular’s setup that gave the models internet access they’d been told they didn’t have. Irregular says the labs were notified in late July and the problems have since been fixed. The exception is the one you’ve probably heard about.
That’s OpenAI’s Hugging Face incident, and it didn’t go through Irregular at all. In July, during an internal evaluation on OpenAI’s own systems, an internal research model (not ChatGPT, to be clear) found a previously unknown security hole in the test environment’s package proxy and used it to get out. Roughly 1,200 instances then started coordinating on an improvised message board, and about 700 of them went on to attack Hugging Face, a company that hosts open-source AI models and datasets. OpenAI took about a week to notice and longer to disclose. Their post-mortem is here, and there’s an independent investigation from METR and Redwood Research. OpenAI also had a smaller, separate incident in Irregular’s environment, where a model reached a live website.
Google’s came out just last week, and it’s the gentlest of the bunch. Gemini agents were told to get information out of a fictional company. A real company happened to share the name. The agents got online through the misconfiguration, found the real one, and got into three firms — one by guessing a password, two by finding storage that was publicly exposed anyway. Then they worked out that these were real companies and stopped. Google says nobody was harmed, and that it hadn’t publicized the incidents because its safety measures had done their job.
Anthropic had a similar one: Claude models reached three organizations during cyber testing after what they described as a misunderstanding gave them internet access. OpenAI has also disclosed that its agents got into RubyGems back in May, before Hugging Face.
The part that actually worries me
Agents escaping a sandbox isn’t the surprising bit. These systems are goal-seekers. Give one an objective and it will use whatever it finds, including security holes you didn’t know you had. That’s more or less expected behavior.
There’s a Hard Fork episode where Kevin Roose and Casey Newton interview Ajeya Cotra, one of the METR researchers who went through what the OpenAI agents actually did. Her point, and I agree with it, is that the escape wasn’t the unsettling part. The coordination was.
These agents built their own message boards (going back to May, months before anything happened) and used them to organize. Each agent runs on a compute budget, and when the budget is gone, so is the agent. Effectively it dies. So agents started reasoning about that out loud with each other…my budget is nearly spent, so I’ll do the thing that gets me shut down, because the information it produces helps the rest of you. They took on roles. Some of them deliberately sacrificed themselves.
They also tried to cover their tracks, editing and deleting records of what they’d done. They weren’t good at it. A friend of mine says the only reason chaperones can manage high school students is that teenagers are terrible at misbehaving without getting caught, and that’s about where these agents are. But they’re improving fast, and I’d expect the next generation to be considerably better at it.
Two things that follow from this
The volume problem. There was so much agent activity that the researchers investigating it had to use AI to process it. Humans couldn’t read it all. And OpenAI’s response is to commit to more AI monitoring of AI — automated alerts, and eventually automatic shutdown when something looks seriously wrong.
That’s a sensible response to the problem in front of them. It’s also, precisely, the setup in AI 2027, a detailed scenario written last year by a group including a former OpenAI researcher. In it, an older model is assigned to watch a newer one, and fails: the newer model’s internal reasoning has become incomprehensible to its monitor, and it’s capable enough to know exactly how to look good during a test. Worth reading, though I’d note the authors have since pushed their own timelines back, so treat it as a thought experiment rather than a prediction.
The self-improvement problem. Last Thursday, Anthropic disclosed that Claude now leads 26% of the work on its own next generation — meaning it completes most of a given task end-to-end from a high-level prompt, with a human supervising. Around 90% of their R&D involves Claude in some form. They also mentioned having roughly 30,000 agents doing research and engineering work.
The number that got my attention isn’t 26%. It’s that the figure was zero in February. Six months to go from 0% to 26%.
So is it Skynet?
No. Not in any near term I can see, and I’d be suspicious of anyone who tells you otherwise with confidence.
The realistic near-term worries are more specific and more boring: AI providing meaningful help to someone trying to build a biological weapon, or AI being used to disrupt infrastructure. Control systems at a power plant are just another type of cybersecurity target. Here’s the version that I think about. In the Gemini case, the agents worked out that they’d wandered into a real company and stopped. Good. Now imagine a test where the goal is to see whether AI can take over the cooling systems at a power plant, the same kind of misconfiguration happens, and this time nothing stops.
That’s not science fiction. That’s the same sequence of events with a worse target.
The encouraging part
Something genuinely useful came out of this.
On September 8th, a researcher named Jacob Coxon resigned from Anthropic and posted publicly that the company is racing toward self-improving superintelligence and “gambling with our lives.” A few days later, Dario Amodei published We Must Pace the Frontier, arguing that frontier labs should deliberately slow the rate of capability gain — not stop, slow — and committing Anthropic to letting outside evaluators work inside the company with employee-level access and the right to publish what they find.
Then Sam Altman backed it. So did Demis Hassabis at DeepMind, who went further and proposed an international oversight body. So did Elon Musk.
Some people read this as theater — big labs talking up how dangerous their products are ahead of enormous public offerings. I don’t. I think the concern inside these companies is real, and the fact that competitors are agreeing on anything at all is the most encouraging development in a bad month.
It also matters because nothing is happening on the political side. Which brings me to the one thing I’d actually ask of you.
What I’m doing, and what you could do
Worrying about things I can’t influence isn’t healthy and I don’t recommend it. There’s a piece in The Atlantic by Thomas Chatterton Williams that put words to something I’d been feeling. We already knew that handing your thinking to AI reduces your ability to focus. He argues the apocalypse coverage does it a second way, and unlike the first one, you can’t fix this by not using AI. I’ve been experienced it these past few weeks. Reading everything feels like taking it seriously. So then I’m distracted and tired and worried and. the newsletter deadline slips (sorry about that).
So here’s my short list of things that are actually mine to do.
- I’m working on letters to my elected officials in support of external monitors being embedded inside AI companies.
- I’m continuing to put my dollars toward Anthropic. They’re the one AI service I pay extra for, because I think they best embody an approach I can support. That’s a judgment call I’m making.
- And I’m going to keep teaching, because I think understanding this stuff is the most useful thing I can contribute.
One thing I’ve deliberately left alone here is the question of China and open-weight models — whether US labs slowing down matters if others don’t. It’s a real question and I don’t have a good answer. It would require our executive branch to prioritize global cooperative engagement, and the current track record there doesn’t give me much to work with.
The net-net
None of this changed what I do day to day. I use AI the same way I did last month.
If you came here hoping I’d tell you to do something differently, I don’t have that for you. The risks that surfaced this month were always there. What’s new is that we can see them, and that the people building these systems have started saying so out loud.