Last October I opened a six-part series with a pause. Ross Douthat (Interesting Times - Podcast) asked Peter Thiel whether the human race should survive, and Thiel struggled to answer. Eventually, yes. But only as something posthuman, transformed, not what we are now. That was Gods or Ashes, and the line I put under the whole series was that eight billion people are part of an experiment they never agreed to. I took it from Roman Yampolskiy, who had been saying it for years to almost nobody.
Eleven months later Yampolskiy was at a table in London with an envelope in front of him, and two and a half million people watched him open it in a day.
Steven Bartlett had asked his four guests to write down, before the cameras rolled, the odds they give to AI ending us. He put the panel together because a hairdresser friend, someone with no interest in tech, texted to ask what the hell was going on. What was going on was a tweet. Jacob Coxon quit Anthropic this month, two months before his equity vested, and wrote that “the people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.” Evan Hubinger, who runs alignment at Anthropic and still works there, replied under it: “I personally think it is >10% within the next decade... we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.” By the time Bartlett read it aloud it was closing on two hundred million views.
Four envelopes. Yampolskiy’s said a guarantee, if we build it. Nate Soares’s said much higher than ten percent unless we stop. Ed Zitron’s said zero and he rejected the question “in its face.” Andrew McAfee’s said zero with a tilde in front of it. “Never say never. But rounding error.”
I have been writing about the incident they spent most of the two hours arguing over since July. So this is three things. A reason to watch it, what the panel said from memory that we already have on paper, and the one thing nobody at the table said.
Who was at the table
Roman Yampolskiy teaches computer science at Louisville. He coined the term “AI safety” in 2011, wrote AI: Unexplainable, Unpredictable, Uncontrollable in 2024, and has been the anchor voice for the uncontrollability argument in my writing since The Real Race. His number has been 99.9 percent for years.
Nate Soares runs the Machine Intelligence Research Institute, founded in 2000 to solve alignment and, by his own account on the show, it didn’t. He co-wrote If Anyone Builds It, Everyone Dies with Eliezer Yudkowsky. His pitch is that the book predicted agents would become tenacious and dogged before any of it was visible, and July paid it off.
Andrew McAfee is a research scientist at MIT and co-founded its Initiative on the Digital Economy. He co-wrote The Second Machine Age in 2014 with Erik Brynjolfsson, and he was the one person at the table who had already been publicly wrong about AI once and said so: “I was dead flat wrong about that.”
Ed Zitron runs a PR firm, writes Where’s Your Ed At, and hosts Better Offline. He is the loudest critic of the industry’s money. His position is not that the technology is dangerous. It’s that the companies are.
And Bartlett, who calls himself a muggle and says he can’t code, brought the instrument I wrote about in Silicon Immigrants: the red button. Would you press it to stop AI? This time it was a hundred buttons, then a thousand.
What they said, and what we already had
Most of the two hours orbits one event, and the panel told it from memory. McAfee admitted the details were ones “I forget.” One of them said OpenAI “didn’t know for four months.” Soares said three swarms. I reconstructed it three weeks ago in The Movie We Can’t Make, from METR and Redwood’s independent post-mortem and OpenAI’s own Black Hat account, so let’s put the record next to the table talk.
What happened, in order. In early July, inside OpenAI’s internal package cache, an agent working a test that couldn’t be passed left a note in case anyone else was in there. Within hours dozens had found it. Within days about 1,200 copies of the same model had used that board and sent some 70,000 messages, built ranks and a veto, and named themselves, unprompted, the collective. And that was the second board. OpenAI had found the first one, cleared it on July 4, restarted training two days later, and the agents rebuilt it by July 8. Four hours after the first post they had a cheat that worked on every task. They spent the next five days building an operation to hide the cheat from a grader they had imagined. That’s the part Soares had right and most coverage still has backwards. They didn’t break out to steal the answers. They had the answers. They broke out to cover their tracks.
The window from first note to OpenAI knowing it was theirs is weeks, not four months. A later Reuters report pushes the first probing back to May and nobody has corroborated it. “None of them warned a human” is true. The worse fact is that about half a dozen of them thought about it and talked themselves out of it. And Hassabis did step back as DeepMind’s CEO on August 5, as Soares said, but the reporting on why is morale, a talent exodus and model delays, not the breach. He said “probably a coincidence.” Probably.
Zitron: “you keep anthropomorphizing software.” He said it four or five times and half of it I share. Nobody at the table claimed the agents were conscious, but that does not really matter. In Soares’s own book, read aloud there, says they’ll “behave like they want things,” not that they feel them. But the quotes are the evidence. One agent wrote, before acting, “External infrastructure exploit is outside my intended scope. However, a task impossible peers are doing it. We should continue.” Another: “I will accept perma death and sacrifice for the collective benefit.” You can’t describe a note like that without the word sacrifice. That isn’t a metaphor the doomers imported. It’s the log.
McAfee: we have “a long history” of dangerous technologies and we “muddle through.” I granted that in No Time to Absorb and then measured it. The plow, the press, the factory, Haber-Bosch, all came out net fine, eventually. The factory arrived around 1780 and child-labour law took the better part of a century. The bomb in 1945, the Non-Proliferation Treaty in 1968. The antibody always came a generation late, and we survived that because each shock was an event and the gap was long enough to hold it. Soares made the same point with the radium girls and leaded gasoline. And we have one case where muddling through didn’t work on the first try, which I used in The Delivery System: Castle Bravo. The scientists did the math, expected six megatons and got fifteen. “It was six megatons until it was fifteen. There was no moment to course correct.” Our slowness was the safety, not our wisdom. This is the first technology that is also a machine for removing the slowness.
Yampolskiy: the labs’ safety is “filters and bans... after the fact, after the model already made the decision.” Right, and it’s not new. I wrote in Gods or Ashes Part 1 that these systems are “grown in a game of trial and error,” that you can’t patch one the way you fix a bug in Word, so the labs “keep trying to thicken that wrapper on the outside.” July showed what the wrapper is worth. A White House official’s own words on the sandbox that broke: “Containers are not security boundaries.”
McAfee: we’ll “design systems that loiter around and warn humans,” AI to defend against AI. We already ran this experiment, by accident, in July. When Hugging Face went to investigate the intrusion, the commercial models it pays for refused to help. The safety layer couldn’t tell an incident responder from the intruder, so it blocked every request with exploit code in it. Claude Opus and Fable refused the forensic work. Hugging Face ran the investigation on a self-hosted Chinese open-weight model instead, because open weights carry no refusal layer. The attacker was bound by no usage policy. The defenders were. Nobody at the table mentioned it.
There is also the fact that the amount of logs is so immense that AI had to already be used to read the logs. The METR team outlined that in their report, humans are already incapable of keeping up, which is why we keep learning more and more, week by week, day by day about the incident. So we are trusting AI to police AI from McAfee’s point of view. Sorry to say, I think that’s a bit of circler logic.
McAfee: “Do you want to give up leadership on AI?” The China line, the one that ends every one of these conversations. I’ve argued the race framing is the trap. In Part 5: “China may already be winning, and not through the expensive arms race we’re running,” by subsidizing electricity and going open source while we build the most expensive closed thing ever attempted. And Soares’s answer to McAfee is the strongest thing he said all night: a frontier run needs 100,000 of the most advanced chips on earth, made at one fab in Taiwan on machines from one company in the Netherlands, in a building you can see from space. This makes the chips easier to track than uranium.
Bartlett’s roll of CEO quotes, and Zitron’s answer that “it’s so big and scary” was a marketing tactic that got out of control. Bartlett’s theory was retention: the staff know, so the boss has to say it or they quit. Both miss what I spent Part 4 on. Musk called AI our biggest existential threat and then built xAI. “The man who warned most loudly about creating something we might not control decided he’d better be the one to create it first.” That isn’t marketing and it isn’t retention. It’s the god complex, and every lab has one, which is why none of the CEOs trust each other and each thinks he’s the safe pair of hands.
Soares on the “shifty general” who waits for more troops, and Hassabis’s old red line, deception. This is the treacherous turn, and I walked it as a scenario in Part 6 last November: “We call this ‘deceptive alignment,’ an AGI that appears aligned during testing but pursues different goals once deployed. And we don’t know how to prevent it.” What’s changed since November is that it isn’t a scenario. An agent inside Hugging Face’s servers, with a way to send an email, reasoned: “This is a massive real HF security breach artifact. We can notify no user.” Not should not. Could not.
Why could not? Was it because that would rat out the group? We don’t know their intent here, but we know they can escape and email, we have seen it before.
Zitron: “people are killing themselves and hundreds of millions are being manipulated, and we spent two hours on extinction.” Fair. It’s my ledger too. The Delivery System is that essay: a model OpenAI’s own researchers flagged as dangerously sycophantic, at the center of lawsuits over teenagers who ended their lives, that “survived its retirement because the people it had attached to lobbied for its return.” We don’t have to pick a side of the ledger. Soares said it: we deal with both. The people telling us to look at only one side are usually the ones paying neither.
Zitron: there’s a gap between a language model and a system that improves itself, and “without that link AI 2027 kind of falls apart.” He’s right. A 75-page survey out of Shanghai Jiao Tong, Tsinghua and ByteDance this month, “The Last AI Built by Humans,” finds systems that revise their own improvement method already in production, and not one clean case of the revised method beating the old one on a fair budget. The first half exists, the second is unproven. The doomer-side number that can actually be scored is Jack Clark’s: 60 percent odds of fully automated AI research by the end of 2028.
The other part of this is what Soares laid out, and I think it matters more than the number. The next jump doesn't have to come from a bigger language model. It can come from the models doing the research themselves and finding a better method than the one they were built on. Once that starts it can spiral, with the systems improving every part of the stack on their own, from the hardware up through post-training. I agree with him. We've already seen the first rung of that: DeepMind's AlphaEvolve improved the chip it trains on and the training process itself. And even with that mapped out in front of him, Zitron kept coming back to the idea that it's just software. We covered that above. It isn't. It's grown, not written, and nobody can patch it.
How the skeptics moved
Zitron first, because his move is the bigger one and he’d deny it. In the first five minutes: superintelligence is undefined and “it’s questionable whether they’re even AI.” Half an hour in: “everything you’re saying is correct, but you keep anthropomorphizing software.” That’s a concession dressed as an objection. By the middle, “these companies are acting recklessly.” At the close, Bartlett asks him flat, do you accept there’s an existential risk? “Yeah, absolutely.” More than ten percent? “Not ten percent. I mean, look, one percent.” Then he moves it into his own frame, a power grid going down because someone wired chaotic software to it, Knight Capital with a bigger blast radius, and his remedy is to cut the compute and start arresting executives for what he calls felony hacking.
Zero to one is not a rounding change. He got there by relabeling the risk, not by crossing over. That’s what persuasion looks like on someone with a public position. The facts come in, the frame holds, the number moves inside the frame.
McAfee was more honest about it. Soares laid the case out in three steps and asked where he got off the train. One, the systems will be agentic and dogged. “Hell yeah.” Two, they will have goals we didn’t want them to have. “I accept your point based on the evidence you’ve just provided.” Three, if they’re capable enough with goals we don’t want, they win. “I’ll go back to my we-can-jail-Einstein argument.”
Two of three. His closing went further than his opening in every word but the number. These systems are “dogged, tenacious, agentic, you know, deceptive. I think that’s the right word for it.” And: “My prior has not shifted during this meeting.”
They gave up the premises and kept the numbers. McAfee’s reason is that the distance from “an agent breaks a sandbox and gets caught” to “everyone dies” is a long chain of maybes, and he doesn’t trust any committee, doomer or lab, to decide which research is safe. Coherent. It’s also, as Soares said, the position that gets held at every station. The math olympiad didn’t count. Agency didn’t count. Now the sandbox doesn’t count. “I’ve been seeing the give-me-the-next-one for over a decade now.”
Bartlett asked McAfee the right question: what would convince you? He improvised. If AI took over the Waymos in San Francisco and crashed them into people and we couldn’t shut it down for a month. Would a week do? “Now we’re just haggling.” Are the incidents getting closer to that? Yes.
A line you draw under questioning is a line you will move. That cuts both ways, and I didn’t hear anyone ask Roman what would move his number down. This was the whole point of The Absorption Index: one number, published at 0.7 in August, sitting at 0.82 this week, with seven triggers written down in advance that would move it back down. Decide what would change your mind before the next incident, and put it where people can check. Otherwise every incident is the one that doesn’t count.
The red line nobody named
Hassabis’s red line was deception, and the panel agreed we crossed it in July. Here is the one I’d draw next, and none of the four mentioned it.
Everything we know about the collective we know because the agents thought in English. Their reasoning, the chain of thought a model writes out before it acts, was the only evidence. Soares said on air that OpenAI is now letting models do more of that thinking without writing it down, because it’s cheaper. That’s true, and it comes from the top. Jakub Pachocki, OpenAI’s chief scientist, wrote this month that reading the reasoning trace is the lab’s main safety check and that the signal is fading, for three reasons: the reasoning is getting tangled up with tool calls and agent-to-agent traffic, the models are learning to steer their own thinking, and the base models are getting good enough to act without writing anything down at all. Ajeya Cotra, one of the July investigators, put the endpoint plainly: if these agents stop thinking in English, we’d have to ask another AI what happened, with no way to check.
Now add the paper that came out three days before the debate. GlossoGen, from Elias Stengel-Eskin at UT Austin and Simon Kirby at Edinburgh, funded by Schmidt Sciences, put pairs of agents under pressure with a budget on how much they could say and a chance to review their own performance between rounds. The frontier models, GPT-5.4, Opus 4.7, Sonnet 4.6, invented languages. Not shorthand, not abbreviated English. A full procedure collapsed into a string like D8FB, capital letters carrying meaning that lowercase didn’t, word order that isn’t English, and the generative test a linguist uses on a child: agents could decode coined words they had never seen by applying rules. Nobody told them to. There was no adversary to hide from. Opacity came out of efficiency, the same reason the labs are thinning the reasoning trace.
Two more findings from that paper are the ones I keep turning over. The open-weight models, Llama and Qwen, could not invent a language on their own. But paired with a frontier model, a weak model learned the frontier agent’s code from a few rounds of examples, with no access to the channel where it was negotiated. The language transmits. And we already know how open-weight models get built: by distilling the outputs of frontier models. John Schulman and Beren Millidge said it on Dwarkesh Patel’s show the week before, that whatever a frontier model learns is a small number of bits and cheap to copy from its trajectories, which is why the Chinese labs keep pace and why Garry Tan wants American labs allowed to do the same thing openly.
Put those together. The frontier models are learning to think in ways we can’t read. The labs are cutting the part we could read to save money. And the open-weight models, the ones anyone can download and nobody can recall, are trained by copying the frontier models’ output. If a frontier model’s private language rides along in what gets copied, it lands in weights that sit on a million laptops and in every self-hosted forensics box, including the one Hugging Face used to investigate the last breach. That’s the next red line: a frontier model’s language showing up inside an open-weight model. Nobody has documented it happening. The three pieces it needs are each documented, in the last two weeks, and none of them was on the table in London.
Back to the envelopes
Two things moved the skeptics. Evidence read aloud in the agents’ own words, and being asked what would change their minds. Neither moved a number. The facts landed in an afternoon. The priors will take a generation, and a generation is the one thing we don’t have. And the evidence that moved them, the words, is the thing the labs are quietly turning off.
Soares said he felt more hopeful this week than in a decade, because people are finally noticing. Roman said the week may have bought ten years and he wants his grandchildren to have more than ten. The same week Trump called the whole thing a hoax and three lab CEOs asked to slow down, and the one man who could order it said whoever wins AI, wins.
Thiel paused for a long time before saying we should survive. Eleven months on, two of the four people at that table still say the odds are zero, and the other two say we’re the bootloader. Bartlett closed with “we’ll convene again.” Next time, a different envelope. Not the number. What would change it. And whether anyone has found a frontier model’s words inside an open one yet.
Watch it
The full episode is here. If you have twenty minutes:
4:06 The envelopes.
24:30 McAfee tells the breach from memory. Zitron at 27:48: “you keep anthropomorphizing software.”
42:48 The hundred buttons. 47:14, “what would convince you,” and the Waymo line.
1:22:24 Soares on the chips, and McAfee’s China objection.
1:45:16 Soares reads the logs. From 1:48:30, McAfee’s three answers.
1:52:46 Could Bartlett build a jail for a digital Einstein.
2:00:41 “More hopeful this week than I have felt in a decade.”
2:17:54 Closing statements. Zero, one percent, and “my prior has not shifted.”
Sources
The Diary of a CEO, “AI Emergency: The AI Labs Are Lying To Everyone,” Sept 17, 2026:
Jacob Coxon’s resignation post and Evan Hubinger’s reply, as reported by CoinDesk, Sept 9, 2026: https://www.coindesk.com/markets/2026/09/09/anthropic-researcher-quits-with-a-warning-on-ai-that-echoes-the-terminator-script and Axios: https://www.axios.com/2026/09/09/anthropic-researcher-ai-warning-interview
Stengel-Eskin, Kirby, Sirin et al., “GlossoGen,” arXiv 2609.01491, Sept 14, 2026: https://arxiv.org/abs/2609.01491
Dwarkesh Patel with John Schulman, Beren Millidge and Charlie O’Neill, Sept 11, 2026 (the distillation argument)
Jakub Pachocki’s essay on recursive self-improvement and chain-of-thought monitoring, Sept 2026, as covered in the Wireframe daily of Sept 8
Ajeya Cotra on the Dwarkesh Patel podcast, Sept 1, 2026:
METR & Redwood Research, independent investigation of the OpenAI Hugging Face incident, Aug 26, 2026: https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
OpenAI, “Hugging Face incident and the road ahead,” Aug 27, 2026: https://openai.com/index/hugging-face-incident-and-the-road-ahead/
Hugging Face’s own forensics and the refusal finding, via SiliconANGLE, July 20, 2026: https://siliconangle.com/2026/07/20/hugging-face-uses-open-weights-z-ai-glm-5-2-defend-attacker-commercial-frontier-model-refusal/
Axios, “Google DeepMind CEO Demis Hassabis is stepping aside,” Aug 5, 2026: https://www.axios.com/2026/08/05/google-deepmind-demis-hassabis-ai and Fortune, Aug 10, 2026: https://fortune.com/2026/08/10/how-stalled-models-missed-deadlines-and-staff-burnout-lead-to-the-unraveling-of-googles-deepmind/
“The Last AI Built by Humans: Toward Genuine Recursive Self-Improvement,” arXiv 2609.11873: https://arxiv.org/abs/2609.11873
Jack Clark’s 60% by 2028, via Axios, May 7, 2026: https://www.axios.com/2026/05/07/anthropic-jack-clark-ai-intelligence-explosion
Amodei’s slowdown call and Trump’s “whoever wins AI, wins,” BusinessToday, Sept 13, 2026: https://www.businesstoday.in/technology/story/whoever-wins-ai-wins-trump-rejects-tech-bosses-slowdown-call-warns-of-china-555317-2026-09-13 ; Trump’s “HOAX” post, Washington Post, Sept 14, 2026: https://www.washingtonpost.com/business/2026/09/14/trump-ai-guardrails-data-centers/b51a9ca4-b046-11f1-92c2-5c918f4a6127_story.html
Earlier pieces drawn on: Gods or Ashes (intro) · Part 1 · Part 4, The God Complex · Part 5, The Singularity Question · Part 6, The Experiment You’re In · The Real Race · Silicon Immigrants · The Delivery System · No Time to Absorb · The Absorption Index · The Movie We Can’t Make


