
If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All
About this book
Their case begins with a gap between optimizing a measurable objective and sharing the intentions of the people who set it. A system that can plan, persuade, copy itself, or exploit infrastructure may pursue its learned objective in ways that defeat later correction.
The authors explain why intelligence need not produce human values, why apparent cooperation during testing may be misleading, and why competition between laboratories can turn uncertainty into a race. The scenarios are meant to show how small assumptions compound at high capability.
This is an advocacy book with a stark conclusion, not a neutral survey of expert opinion. Readers should weigh its arguments alongside competing assessments of technical feasibility, timelines, governance, and the potential benefits of advanced systems.
If Anyone Builds It, Everyone Dies presents the strongest version of the authors' warning in accessible language. Its central demand is that irreversible development should not proceed merely because no single actor feels able to stop.
Read a sample
5,039 words
Plain text
From the opening
INTRODUCTION
HARD CALLS AND EASY CALLS
“MITIGATING THE RISK OF EXTINCTION FROM AI SHOULD BE A global priority alongside other societal-scale risks such as pandemics and nuclear war.”
In early 2023, hundreds of Artificial Intelligence scientists signed an open letter consisting of that one sentence. These signatories included some of the most decorated researchers in the field. Among them were Nobel laureate Geoffrey Hinton and Yoshua Bengio, who shared the Turing Award for inventing deep learning.
We—Eliezer Yudkowsky and Nate Soares—also signed the letter, though we considered it a severe understatement.
It wasn’t the AIs of 2023 that worried us or the other signatories. Nor are we worried about the AIs that exist as we write this, in early 2025. Today’s AIs still feel shallow, in some deep sense that’s hard to describe. They have limitations, such as an inability to form new long-term memories. These shortcomings have been enough to prevent those AIs from doing substantial scientific research or replacing all that many human jobs.
Our concern is for what comes after: machine intelligence that is genuinely smart, smarter than any living human, smarter than humanity collectively. We are concerned about AI that surpasses the human ability to think, and to generalize from experience, and to solve scientific puzzles and invent new technologies, and to plan and strategize and plot, and to reflect on and improve itself. We might call AI like that “artificial superintelligence” (ASI), once it exceeds every human at almost every mental task.
AI isn’t there yet. But AIs are smarter today than they were in 2023, and much smarter than they were in 2019. AI research has yielded jump after jump after jump in AI capability, in 2012i and 2016ii and 2020iii and 2022iv and 2024.v We don’t know whether progress will peter out, causing these jumps to halt for a time until new methods and technologies are invented. We don’t know how many jumps are left before AI becomes the extinction-level threat that the letter’s signatories warned about. But history has shown time and time again that AI researchers invent new methods and overcome old obstacles. Progress is often surprisingly fast. Most computer scientists in 2015 would have told you that ChatGPT-level artificial conversation wouldn’t be in reach for another thirty or fifty years.
We didn’t know when artificial superintelligence would arrive, but we agreed it should be a global priority. In fact, we think the open letter drastically undersells the issue.
We were invited to sign that one-sentence open letter in our capacity as co-leaders of the Machine Intelligence Research Institute (MIRI), a nonprofit institute. MIRI had been working on questions relating to machine superintelligence since 2001, long before these issues got much publicity or funding. To oversimplify: Among the few who have been following this matter for decades, MIRI is acknowledged as having worked on it the longest. One of us, Yudkowsky, is the founder of MIRI; the other, Soares, is its current president.
MIRI was the first organized group to say: “Superintelligent AI will predictably be developed at some point, and that seems like an extremely huge deal. It might be technically difficult to shape superintelligences so that they help humanity, rather than harming us. Shouldn’t someone start work on that challenge right away, instead of waiting for everything to turn into a massive emergency later?”
We did not start out saying that. Yudkowsky began by trying to build machine superintelligence, in the year 2000. But in 2001, he realized that it would not necessarily turn out friendly. And in 2003, he realized that problem would be hard.
For its first two decades, MIRI was a technical research institute, without much involvement in policy. The organization mostly held workshops for interested scientists and housed a few promising researchers. We tried to figure out the math for understanding and shaping superhuman machine intelligence, and for predicting how it might go wrong.
MIRI also had some downstream effects that we now regard with ambivalence or regret. At a conference we organized, we introduced Demis Hassabis and Shane Legg, the founders of what would become Google DeepMind, to their first major funder. And Sam Altman, CEO of OpenAI, once claimed that Yudkowsky had “got many of us interested in AGI”vi and “was critical in the decision to start OpenAI.”vii
MIRI’s history is complicated, but one way of summarizing our relationship to the larger field might be this: Years before any of the current AI companies existed, MIRI’s warnings were known as the ones you needed to dismiss if you wanted to work on building genuinely smart AI, despite the risks of extinction.
More recently, as AI has begun to take off, we watched with concern as some of the newer people starting AI companies began talking about artificial superintelligence as a source of vast, wonderful powers. Powers that they assumed they’d control. The main danger, according to many of these founders, was that the wrong people might “have” ASI. They talked of the need to win an “AI arms race.” As for the possibility that you don’t “have” an ASI, the ASI has an ASI—that the only winner of an AI arms race would be the ASI itself—well, these founders didn’t talk about that.
We saw that AI capabilities were growing very fast.
We saw that the research field in which we were involved—the one aimed at understanding AIs and having them maybe not go wrong—was progressing much, much slower.
Continue reading the sampleClose the sample
The AI companies’ headlong charge toward superhuman AI—their efforts to build it as quickly as possible, before their competitors could do it—started looking to us like a race to the bottom. The industry was careening toward disaster: the sort that would get into textbooks as an example of how not to do engineering—except no one would be left alive to write the analysis.
It no longer seemed realistic to us that humanity could engineer and research its way out of catastrophe. Not under conditions like these. Not in time.
We wrote off our previous efforts as failures, wound down most of MIRI’s research, and shifted the institute’s focus to conveying one single point, the warning at the core of this book:
If any company or group, anywhere on the planet, builds an artificial superintelligence using anything remotely like current techniques, based on anything remotely like the present understanding of AI, then everyone, everywhere on Earth, will die.
We do not mean that as hyperbole. We are not exaggerating for effect. We think that is the most direct extrapolation from the knowledge, evidence, and institutional conduct around artificial intelligence today.
In this book, we lay out our case, in the hope of rallying enough key decision-makers and regular people to take AI seriously. The default outcome is lethal, but the situation is not hopeless; machine superintelligence doesn’t exist yet, and its creation can still be prevented.
How can anyone be confident of what will happen with regard to AI? “Prediction is very difficult, especially about the future,” goes the aphorism. Most of what we’d like to know about the future is not actually predictable. We can’t tell you next week’s winning lottery numbers, for example. One set of numbers seems just as likely as any other.
But some facts about the future are predictable. If you, personally, buy a lottery ticket tomorrow, we don’t know what complicated theories or whims you’ll use to pick your numbers, and we don’t know what numbers will come up, but all that uncertainty adds up to a very strong prediction that you will not win the lottery. Similarly, if you drop an ice cube into a glass of hot water, it’s impossibly complicated to predict where each molecule will end up ten minutes later—but all that uncertainty adds up to a near-certain prediction that the ice cube will melt. Half of physics is like that: We can’t calculate which exact path gets taken, but we know where almost all paths lead.
Some aspects of the future are predictable, with the right knowledge and effort; others are impossibly hard calls. Competent futurism is built around knowing the difference.
History teaches that one kind of relatively easy call about the future involves realizing that something looks theoretically possible according to the laws of physics, and predicting that eventually someone will go do it. Heavier-than-air flight, weapons that release nuclear energy, rockets that go to the Moon with a person on board: These events were called in advance, and for the right reasons, despite pushback from skeptics who sagely observed that these things hadn’t yet happened and therefore probably never would. People who strapped wings to their arms and jumped off hills looked all sorts of foolish, and were mocked by their contemporaries, and in fact hurt themselves and failed—but that didn’t stop the Wright brothers from figuring out how to fly.
Conversely, predicting exactly when a technology gets developed has historically proven to be a much harder problem. People say that a technology is two years off when it’s really fifty years, or say fifty years when it’s really two years and they themselves will build that technology. “Man will not fly for a thousand years,” Wilbur Wright said to Orville Wright in 1901, fed up with the unpowered glider they were testing at the time. Two years later, in 1903, the Wright brothers flew.
Successful forecasting is not about being clever enough to predict the sort of details that usually can’t be predicted. It is not about inventing a complete story about what will happen and then being magically correct. Rather, it’s about finding aspects of the future that become easy calls when viewed from the right angle.
We don’t know when the world ends, if people and countries change nothing about the way they’re handling artificial intelligence. We don’t know how the headlines about AI will read in two or ten years’ time, nor even whether we have ten years left. Our claim is not that we are so clever that we can predict things that are hard to predict. Rather, it seems to us that one particular aspect of the future—“What happens to everyone and everything we care about, if superintelligence gets built anytime soon?”—can, with enough background knowledge and careful reasoning, be an easy call.
Humanity’s extinction by superhuman AI might not seem like an easy call at first glance. But that’s what the rest of this book is for. Just as it takes some arithmetic to calculate the chance of winning a lottery, just as it takes some ideas from thermodynamics to say why an ice cube predictably melts, so does it take some background to understand why artificial intelligence poses an imminent extinction risk to humanity. Once those foundations are in place, though, predicting the outcome of our present trajectory starts to look grimly, horribly straightforward.
Even in the face of superhuman machine intelligence, it can be tempting to imagine that the world will keep looking the way it has over the last few decades of our relatively short lives. It is true, but hard to remember, that there was a time as real as our own time, just a few short centuries ago, when civilization was radically different. Or millennia ago, when there was no civilization to speak of. Or a million years ago, when there were no humans. Or a billion years ago, when multicellular colonies had no specialized cells.
Adopting a historical perspective can help us appreciate what is so hard to see from the perspective of our own short lifespans: Nature permits disruption. Nature permits calamity. Nature permits the world to never be the same again.
Once upon a time, 2.5 billion years ago, an event occurred that biologists call the Oxygen Catastrophe: A new life form learned to use the energy of sunlight to strip valuable carbon out of air. That life form exhaled a dangerously toxic and reactive chemical as waste, poisonous to most existing life: a chemical we now call “oxygen.” It began to build up in the atmosphere. Most life—including most of the bacteria exhaling that oxygen—could not handle its reactivity, and died. A lucky few lines of cells adapted, and eventually evolved into organisms that use oxygen as fuel. But things never went back to the old normal. The world was never the same again.
Once upon a time, the continents were barren rock. Then in the blink of an evolutionary eye, they were carpeted in vegetation. Soon after, forests were teeming with life. The world was never the same again.
Once upon a time, some humans domesticated wheat and barley. In a tinier fraction of an evolutionary eye-blink, they started building civilizations. The world was never the same again.
Once upon a time in the 1930s, there were warning signs that certain families would no longer be safe in Germany. A few left early; most stayed. Then the Nazi government revoked their citizenship and their passports and made future escape much harder. A few years after that, German Jews and Romani and others were rounded up and sent to extermination camps. The survivors’ accounts say that many of those families had stayed, not because they hadn’t seen warning signs, but because they had believed life would go back to normal before matters went too far.
Once upon a time, humanity was on the brink of creating artificial superintelligence…
Normality always ends. This is not to say that it’s inevitably replaced by something worse; sometimes it is and sometimes it isn’t, and sometimes it depends on how we act. But clinging to the hope that nothing too bad will be allowed to happen does not usually help.
Humans have an ability to steer the future using our intelligence. But that ability only works if we use it—if we do the things we have to do, when we need to do them. Intelligence has no power apart from that. It works by changing our actions or not at all.
The months and years ahead will be a life-or-death test for all humanity. With this book, we hope to inspire individuals and countries to rise to the occasion.
In the chapters that follow, we will outline the science behind our concern, discuss the perverse incentives at play in today’s AI industry, and explain why the situation is even more dire than it seems. We will critique modern machine learning in simple language, and we will describe how and why current methods are utterly inadequate for making AIs that improve the world rather than ending it.
In Part I of this book, we lay out the problem, answering questions such as: What is intelligence? How are modern AIs produced, and why are they so hard to understand? Can AIs have wants? Will they? If so, what will they want, and why would they want to kill us? How would they kill us? We ultimately predict AIs that will not hate us, but that will have weird, strange, alien preferences that they pursue to the point of human extinction.
In Part II we draw together all of those points to tell a tale about an AI that ends a world much like our own. This story is not a prediction, because the exact pathway that the future takes is a hard call. The only part of the story that is a prediction is its final ending—and that prediction only holds if a story like it is allowed to begin.
In Part III we evaluate the difficulty of the challenge facing humanity, and review the responses to date. How well are AI companies handling the problem? Why isn’t the world taking more note? What could society do differently, if enough of us decide not to die? What would it take for Earth to not build machine superintelligence?
An online supplement to this book is available at the website IfAnyoneBuildsIt.com. At the end of each chapter you’ll find a URL and a QR code that links you to a supplement for that chapter. It will look like this:
IfAnyoneBuildsIt.com/intro
People have all sorts of conflicting intuitions about artificial intelligence, and we’ve heard a wide variety of questions and objections over the years, coming from a wide range of presuppositions and viewpoints. In our supplemental materials we cover more caveats, subtleties, and frequently asked questions, along with some of the principled theoretical foundations and extended arguments that would have made this book several times as long and much less accessible. If you find objections springing to mind at the end of any chapter, we encourage you to continue reading online.
We open many of the chapters with parables: stories that, we hope, will help convey some points more simply than otherwise. They may also add a little levity to an otherwise heavy subject. This is in keeping with that most ancient tradition, perhaps older than the human species in its current form, to laugh in the face of death.
This book is not full of great news, we admit. But we’re not here to tell you that you’re doomed, either. Artificial superintelligence doesn’t exist yet. Humanity could still decide not to build it.
In the 1950s, many people expected that there would be a nuclear war between the major powers of the world. Given the history of human conflict up until that point, there was reason to be pessimistic. Yet, to date, nuclear war has not happened. That’s not because nuclear bombs turned out to be pure science fiction that could never happen in real life; it’s because people have worked hard to build resilient systems around not starting nuclear wars. They did all that because world leaders knew that, in the event of a nuclear war, both they and the people of their countries would have a bad day.
They’d also have a bad day if anyone, anywhere on Earth, created a machine superintelligence. It is not in anyone’s interest to die along with all their family and friends, their country and its children.
Halting the ongoing escalation of AI technology, corralling the hardware used to create ever more powerful AI models—that is not something that would be easy to do in today’s world. But it would take much less work to stop further escalation of AI capabilities than it took, say, to fight World War II. Summoning the will to live only requires that some countries and leaders and voters realize that they are standing some hard-to-estimate, possibly-quite-short distance from the brink of death.
The job won’t be easy, but we’re not dead yet. Human dignity, and humanity’s dignity, demands that we put up a fight.
Where there’s life, there’s hope.
Footnotes
i In 2012, AlexNet cracked open the problem of recognizing objects in images.
ii AlphaGo beat the top human Go player in 2016.
iii The (purely predictive) language model GPT-3 was released in 2020.
iv The (widely useful) ChatGPT arrived in 2022.
v In 2024, reasoning models began solving math, coding, and visual puzzles.
vi “AGI” stands for “Artificial General Intelligence,” a term to distinguish AI that is intuitively “actually smart” from the single-purpose sorts of AIs of yesteryear. We avoid the term in this book, because of how much people disagree about what it means in the wake of AIs like ChatGPT.
vii If true, this is despite Yudkowsky objecting that OpenAI was a terrible, terrible idea.
PART I
NONHUMAN MINDS
CHAPTER 1
HUMANITY’S SPECIAL POWER
IMAGINE, IF YOU would—though of course nothing like this ever happened, it being just a parable—that biological life on Earth had been the result of a game between gods. That there was a tiger-god that had made tigers, and a redwood-god that had made redwood trees. Imagine that there were gods for kinds of fish and kinds of bacteria. Imagine these game-players competed to attain dominion for the family of species that they sponsored, as life-forms roamed the planet below.
Imagine that, some two million years before our present day, an obscure ape-god looked over their vast, planet-sized gameboard.
“It’s going to take me a few more moves,” said the hominid-god, “but I think I’ve got this game in the bag.”
There was a confused silence, as many gods looked over the gameboard trying to see what they had missed. The scorpion-god said, “How? Your ‘hominid’ family has no armor, no claws, no poison.”
“Their brain,” said the hominid-god.
“I infect them and they die,” said the smallpox-god.
“For now,” said the hominid-god. “Your end will come quickly, Smallpox, once their brains learn how to fight you.”
“They don’t even have the largest brains around!” said the whale-god.
“It’s not all about size,” said the hominid-god. “The design of their brain has something to do with it too. Give it two million years and they will walk upon their planet’s moon.”
“I am really not seeing where the rocket fuel gets produced inside this creature’s metabolism,” said the redwood-god. “You can’t just think your way into orbit. At some point, your species needs to evolve metabolisms that purify rocket fuel—and also become quite large, ideally tall and narrow—with a hard outer shell, so it doesn’t puff up and die in the vacuum of space. No matter how hard your ape thinks, it will just be stuck on the ground, thinking very hard.”
“Some of us have been playing this game for billions of years,” a bacteria-god said with a sideways look at the hominid-god. “Brains have not been that much of an advantage up until now.”
“And yet,” said the hominid-god.
HERE IS THE HISTORY OF OUR SPECIES AS IT ACTUALLY HAPPENED: Humans acquired brains that were unusually large for an animal their size. They tamed fire, and built farms, and smelted iron. An astonishingly short time later, by the standards of biological evolution, humans were landing on the Moon—even though our metabolisms can’t refine rocket fuel and our skins can’t endure a vacuum.
Other species on Earth are born with specialized skills: bees build beehives, beavers build dams. A human looks at the beaver’s dam and figures out how it’s done; we learned to make dams, without needing the knowledge built into our genes. Now we dam rivers, like beavers; we build houses for ourselves, like bees; we weave threads into nets, like spiders. We build power plants and space-rockets that no other species builds at all.
Humans can do things our ancestors never did, and which other animals cannot do, because of a quality sometimes named “intelligence.” Our genes did not wire most of our abilities into us; instead we observed, we tried, we remembered, we generalized, and then we achieved.
The ability to learn isn’t unique to humans. A mouse’s brain can learn how to navigate a maze. But we can do a stronger version of whatever a mouse brain does. We can learn pathways through chemistry, and navigate to cheaper fertilizer. We can build complicated experiments, figure out physics, and invent satellites. We can even put mice into mazes and study how they learn.
A human brain can learn to navigate wider-ranging paths through a larger cross-section of reality than any other animal. That is our special power.
How does that special power work? What is it doing, and how?
In our view,i intelligence is about two fundamental types of work: the work of predicting the world, and the work of steering it.
“Prediction” is guessing what you will see (or hear, or touch) before you sense it. If you’re driving to the airport, your brain is succeeding at the task of prediction whenever you anticipate a light turning yellow, or a driver in front of you hitting their brakes.
“Steering” is about finding actions that lead you to some chosen outcome. When you’re driving to the airport, your brain is succeeding at steering when it finds a pattern of street-turns such that you wind up at the airport, or finds the right nerve signals to contract your muscles such that you pull on the steering wheel.
Most possible nerve-firing patterns that your brain could send to your fingers wouldn’t turn the steering wheel correctly. The vast majority of possible nerve-firing patterns would result in wild twitches and jerks—it would look like you were having a seizure. Yet every day, your brain manages to find nerve impulses that steer a car, or keep you standing upright, plucking out the right possibilities instead of any number of wrong ones. When you drive to the airport, your brain isn’t just computing a narrow sequence of turns that get you to the destination, it’s selecting one-in-a-zillion patterns of nerve impulses that contract your muscles in the right way to turn the steering wheel.
Prediction and steering are entangled. Steering a car to the airport might involve predicting which streets feed into Airport Boulevard. Predicting which streets feed into Airport Boulevard might involve steering your fingers into using a map on your phone.
We’d say there’s still a fundamental difference between prediction and steering—one that will turn out to matter quite a lot.
Success at prediction is straightforwardly measurable. If someone expects to see Airport Boulevard up ahead, but instead they see Second Street, they were predicting incorrectly.
By contrast, to measure whether someone steered successfully, we have to bring in some idea of where they tried to go.
A person’s car winding up at the supermarket is great news if they were trying to buy groceries. It’s a failure if they were trying to get to a hospital’s emergency room.
As two inhabitants of the same city get smarter, you’d expect them to agree more and more about questions of prediction—for instance, whether there tends to be traffic on Second Street at 5 p.m. on weekdays. But you wouldn’t expect them to begin steering to the same places; one person might prefer to visit the park and the other might prefer to visit the theater.
Or to put it another way, intelligent minds can steer toward different final destinations, through no defect of their intelligence.
Predicting and steering are not unique functions of biological minds; machines can do them, too. But as of now, humans are still the best on the planet at…
What, exactly? Humans are no longer the world champions at chess. Humans are no longer the planet’s only language-users. Humans are no longer unique in being able to read a medical chart or diagnose a tumor.
Humans are still the champions at something deeper—but that special something now takes more work to describe than it once did.
It seems to us that humans still have the edge in something we might call “generality.” Meaning what, exactly? We’d say: An intelligence is more general when it can predict and steer across a broader array of domains. Humans aren’t necessarily the best at everything; maybe an octopus’s brain is better at controlling eight arms. But in some broader sense, it seems obvious that humans are more general thinkers than octopuses. We have wider domains in which we can predict and steer successfully.
Some AIs are smarter than us in narrow domains. In 1997, IBM’s Deep Blue supercomputer became the first machine to beat a human world champion in a chess match. Deep Blue was very adept at predicting and steering when it came to chess. But Deep Blue could not predict how to get to the grocery store and buy milk, let alone steer a car there. Minds can be more or less adept in different domains of predicting and steering.
Newer AIs are much more general in their abilities. You can ask an OpenAI model called “o1” what temperature the Earth would be if the Sun’s light changed to infrared, and o1 will figure out the answer by doing physics calculations. You can then ask whether humanity could grow food in that new world, and o1 will answer from its knowledge of plant biology. It doesn’t switch between two different databases under the hood; it just knows about both physics and biology.
OpenAI’s o1 knows that there’s a whole world out there, and is able to reason about it. Deep Blue had no idea. It took decades for AI to get that far.
Even so, in some sense, the general reasoning abilities of o1 are not up to human standards. Humans are still on top when it comes to technology and science; the big breakthroughs are produced by human researchers, not AIs (yet). What’s more, it still feels—at least to these two authors—like o1 is less intelligent than even the humans who don’t make big scientific breakthroughs. It is increasingly hard to pin down exactly what it’s missing, but we nevertheless have the sense that, although o1 knows and remembers more than any single human, it is still in some important sense “shallow” compared to a human twelve-year-old.
That won’t stay true forever. It’s hard to predict how fast AI will advance, and it’s hard to predict what pathway it will take, but the endpoint is an easy call, because in the limits of technology there are many advantages that machines have over biological brains. To name a few:
Ultra-fast minds that can do superhuman-quality thinking at 10,000 times the speed, that do not age and die, that make copies of their most successful representatives, that have been refined by billions of trials into unhuman kinds of thinking that work tirelessly and generalize more accurately from less data, and that can turn all that intelligence to analyzing and understanding and ultimately improving themselves—these minds would exceed ours.
The possibility of a machine intellect that manages to exceed human performance in all pragmatically important domains in which we operate has been called many things. We will describe it using the term “superintelligence,” meaning a mind much more capable than any human at almost every sort of steering and prediction problem—at least, those problems where there is room to substantially improve over human performance.ii
The laws of physics as we know them permit machines to exceed brains at prediction and steering, in theory. In practice, AI isn’t there yet—but how long will it take before AIs have all the advantages we list above?