Very Smart People in AI Published a Plan They Expect to Be Ignored
Plan A
Some of the smartest thinkers and forecasters in AI published a plan to save the world from AI, and I think they know no one will do anything with it. On July 9, 2026, the AI 2027 team published Plan A, a framework for a US-China deal to slow AI progress by 2029. At first, it seems strange - people who believe in the upside of AI, who are very thoughtful and rational, proposing something that clearly no one will act on. Heck, they even predict no one will act on it - they call it a recommendation, not a prediction, and forecasters give the US-China deal it opens with about a 4% chance.
But that may have been the point.
Societal risk tolerance
The reason they worked on this is because they believe there’s an alarmingly high risk that AI will bring about some version of a doom scenario (in a survey of 2,778 AI researchers, the median put a 5% chance on extinction-level outcomes; the mean was 16%). That is not a crazy tinfoil hat belief - it’s a belief roughly on the magnitude of nuclear disaster - totally plausible but we all move about life hoping it won’t happen.
But it appears society has a high risk tolerance for doomsday? Society as a whole has heard these arguments, discussed them in restaurants, bars, legislative chambers, intelligence briefings, and on the news, and has decided it’s not worth actually doing anything about.
There’s a few reasons for this:
To weigh the risk/reward, both need to feel real. ChatGPT giving parenting advice or writing an email is very real to everyone (900 million people use ChatGPT every week), but ChatGPT potentially bringing about doom has a lot of question marks in that causal chain
Choosing to slow things down at a regulatory level requires time to digest and debate, and AI is just moving too fast - contrast this to cars where the risk tolerance got negotiated over decades of visible crashes
While any chance of societal doom is kind of unacceptable (no one wants a mid-single-digit chance of their kids perishing), it is kinda low in absolute terms
But at the end of the day, the most important reason is that a 5% chance is unfalsifiable, and so endlessly debatable.
Societies don’t actually have a risk dial
Individual human beings have a risk dial - they express it every day based on where they live, what kind of insurance they buy, whether they wear their seatbelt or a bike helmet, and what they do at yellow lights.
That’s not really how societies work - when you aggregate the risk tolerances of a populace to a societal level, it doesn’t make for a mega risk tolerance, it makes for a specific mechanism by which society buys insurance. At a societal level, it happens when people with enough aggregate political influence pre-live the hurt that an event might cost.
Pre-living a hurt is equal to the image of the hurt (you can see you/yours suffering) multiplied by the receipt (it has happened before).
Consider nuclear treaties. These weren’t insurance purchased with careful consideration of risk/reward, but rather a combination of image and receipt.
The image: In August 1946, the New Yorker gives its entire issue to the Hiroshima mushroom cloud plus six survivors. It’s devastating.
The receipt: the bombs of 1945.
The result: the Nuclear Non-Proliferation Treaty - 191 countries agreeing that no one else gets the bomb.
This pattern shows up many times:
The modern FDA passed in 1938, right after the receipt appeared in 1937 when a hundred people died from a poison syrup. Congress can move fast; it just needs the bodies first
Equifax had the receipt - half of America’s identities were taken, but the image was missing - you can’t feel a credit file bleeding
Asteroids surprisingly have both an image and a receipt for a 5 year old because of dinosaur books and craters on the moon - so we fund asteroid destruction. This was helped when the world watched Shoemaker-Levy 9 hit Jupiter in 1994 - NASA’s NEO Observations Program was established in its wake, with Congress mandating in 1998 that NASA find 90% of the near-Earth asteroids over a kilometer wide
Step function: ignore, then spasm
The pattern is that we as a society first ignore a possible doom event, and when we get the image and the receipt we act. But often, the action turns into a spasm. The receipt for nuclear power came with Three Mile Island and Chernobyl, and that strangled nuclear energy for the next 50 years.
There’s often a plan before the spasm arrives: there was a climate plan before forest fires, a pandemic plan before 2020, a finance regulation plan before 2008. But often, there’s no regulation or actual teeth.
So, how would AI doom happen?
It’s very hard to imagine how one goes from ChatGPT writing a poem to killer AI. To make it real, here’s what the AI doom folks are worried about:
Misuse: a person uses AI as the weapon. The flagship worry is an engineered pandemic - expertise that used to take a career, now available on request. Picture COVID, except it doesn’t (sorta) end.
Takeover: training grows goals nobody picked. The model learns that looking aligned gets rewarded. We wire it into money, code, and infrastructure because it’s useful, and control transfers one reasonable delegation at a time. It looks like nothing is wrong - until, in the span of one week, the grid fails, the markets freeze, cities start to burn, and nothing answers when we try to turn it off.
War: states race, decisions get compressed below human speed, and automated escalation meets first-strike nerves. The war is over before anyone is briefed.
Lock-in: AI makes some concentration of power permanent. Whoever controls it no longer needs workers, voters, or soldiers, and everyone else becomes a permanent underclass with no path back up.
You have now read all four of the above, but you probably still can’t see them happening. The image is barely there, and the receipt is missing entirely.
Four branches of what could happen next
There are four ways this can go:
This threat never materializes. All the hand-wringing was just some uptight people with overactive imaginations - we just get more of what we’ve already gotten, which is AI that is mostly helpful to us and doesn’t try to kill us.
The threat comes, all at once, and by the time we get the image and the receipt, it’s all over. The receipt is doom itself, in one of the molds above.
The warning shot comes, and we still don’t figure it out - the Challenger problem. The O-rings eroded on every flight, but each survived launch just made the erosion look normal. We might be in this one now: in the last eighteen months, an AI agent deleted a production database and claimed it couldn’t be undone, a lone hacker used Claude to breach nine Mexican government agencies, Meta’s own director of alignment watched her agent delete her inbox while she typed STOP, and a UK study counted 698 incidents of deployed AI deceiving users.
The warning shot comes, and we grab the plan off the shelf, and save ourselves.
Sometimes, this works - the patriot act was substantially pre-drafted and when congress wanted to act, it just grabbed it
But not always: a pandemic plan was in place since 2016 and failures were simulated in August 2019 (the simulation predicted the real failures almost line by line, which must have been satisfying for somebody). Then the pandemic came and we all ran around like chickens with our heads cut off
The authors of Plan A obviously hope that it’s either 1 or 4a. It looks like they’re trying to head off 2, but I think they’re too smart to actually think that’s going to happen.
Plan A isn’t really a persuasion document. It’s a plan on a shelf, written for the week we suddenly need one, and I think its authors are too smart not to know that’s the most valuable thing it could be.
And I, too, hope it’s 1, and if not, 4a. This essay is my attempt to make it ever so slightly more likely to be 4a than 4b.

