
An AI safety expert argues that we cannot wait to resolve all uncertainties before pausing AI development.
A former OpenAI, DeepMind, and UK AI Security Institute scientist estimates a 50% chance of human extinction from superintelligent AI within two to ten years, urging an immediate pause on development due to unresolved safety risks like hacking, persuasion, and uninterpretable reasoning.
AI-generated summary
The author worked at OpenAI, DeepMind, and as Chief Scientist at the UK AI Security Institute.
For almost a decade, it's been my job to think about where AI is heading and what could go wrong, first at OpenAI, then DeepMind, then as Chief Scientist at the UK AI Security Institute (AISI). This experience has made it clear to me that we can't wait for all of our uncertainties to be settled to do something about AI's risks.
Recent warnings about the potential destructive power of AI are understating the severity of the situation. I believe thereâs about a 50% chance we all die because of the development of smarter-than-human AI systems, and that our actions over the next two to 10 years will determine the outcome.
To be clear, when I say our odds of being killed by superintelligent AI are 50%, Iâm not claiming precision. Instead, Iâm highlighting that several of the most fundamental debates about the future and safety of AI remain unresolved. I believe we canât expect these disagreements to be resolved until itâs too late to change course.
Ultimately, AIâs capacities and motivations, and our own human actions, all feed into our likelihood of survival. Hereâs why I estimate our odds are comparable to a coin flipâand why I believe the time to pause AI development is now.
The first is hacking. To kill all humans, a superintelligent AI would need to take over and migrate between computer systems. We saw this during the Hugging Face incident. Second, a superintelligent AI would need the ability to persuade, to manipulate humans into acting on its behalf, as we studied at AISI. Third, a lethal superintelligent AI would need to conceal its thoughts from us, which is increasingly common in the latest, most capable models. And lastly, it would need planning and coordination between agents to multiply their effectiveness, like the more than 10,000 agents who worked on the Navier-Stokes problem at OpenAI, or the more than 1,000 involved in the Hugging Face attack.
These skills are very close to those the AI companies intentionally train for. Hacking includes looking for vulnerabilities, just as you would to fix them. Persuasion involves writing text humans like reading. Faster thinking is cheaper, and the pressure to think quickly forces AIs towards using shorthand that is harder for humans to follow. Finally, planning and coordination are useful for tackling ambitious problems in math, programming, or any other domain.
We can be much more confident that a superintelligent AI could take over than about how it would take over, just as I know Lee Sedol would crush me at Go but not which moves he would use to do so. But a variety of broad paths are possible. For instance, AIs could escape onto the internet and gain control of weapons systems. Or AIs could take over the company that trained them and hoard resources.
Letâs imagine the latter. An AI could begin to use concealed reasoning to delay researchers from noticing emerging ill intent because appearing friendly and prioritizing speed help the AI gain reward during training. Later, as the company relies more and more on the AI for a combination of coding help and strategic advice, the AI might use subtle persuasion to reduce resources spent on safety and tamper with experiments designed to detect AI deception. After all, in 2026, when a researcher reads a report about a safety experiment, that report was itself written by an AI.
Once the AIâs hacking skills are sufficient, it could break out of its internal sandbox and sabotage logs of its own misbehavior, as models in the Hugging Face incident attempted to do. Eventually, the AI could become strong enough at hacking and persuasion to move outside the company and take over other data centers, companies, and governments. With this power, AIs might prioritize spending limited resources (such as energy and land) towards their own proliferation, causing life-threatening shortages for humans.
And yet, just because an AI could kill all humans doesnât necessarily mean it will. This question of AI motivation is crucial, but deeply confusing. Different starting points lead different people either to confidence that everything will be fine, or confidence that superintelligence would almost certainly kill us.
Over the last few years, weâve seen many, many examples of model misbehavior at various scales, ranging from blackmail, corporate espionage, and murder in experimental settings to real-world hacks that would likely incur prison time were they committed by a person. None of these involved superintelligence, and we donât know whether these comparatively modest misbehaviors will scale to superintelligent models wanting to kill us all. And âwantâ itself is controversial: people vehemently debate whether itâs sensible to apply anthropomorphic language and thought experiments to AIs at all.
Within the field of AI safety, some researchers think that AI wisdom has been increasing with recent models, and that this may produce wise models that mean well in the future. Others, like Eliezer Yudkowsky and Nate Soares, think the empirical methods this would require are extremely unlikely to work, and place the chance that superintelligence kills us all very high. The disagreement is, in part, about whether the behavior of AIs will be more strongly guided by the values implicitly represented in the human data from early training or by the pressure cooker of later stages of training where AIs are forced to perform better and better at difficult tasks by any means necessary.
I am somewhere in the middle, leaning towards Yudkowskyâs view. But where I land is not the point. Experts in AI safety disagree, but the moderate position is that the risk of extinction is significant. We do not know with confidence whether artificial superintelligence will facilitate human flourishing, or if it will want to kill us to serve its own survival and propagation. But that uncertainty alone is unacceptable.
Other uncertainties surrounding AI matter less. For instance, one of the most persistent debates in AI is the extent to which teaching AIs to do one set of tasks helps them do other tasks, typically referred to as âgeneralization.â A particularly controversial framing is the concept of artificial general intelligence (AGI), âa system that is approximately human level in all cognitive domains.â John McCarthy and Y. Bar-Hillel were debating whether AIs would generalize in 1959. We debated this same topic in 2018 at OpenAI. The whole world seems to be debating it today.
The more AIs generalize from one type of task to another, the faster AI progress will be, and the less time we have until superintelligence arrives. This makes generalization and AGI key concepts when reasoning about the speed of progress. Indeed, people who treat AGI seriously have been much more correct about the speed of AI advances than people who dismiss the concept.
But itâs easy to conflate âAGI, the sometimes-useful concept" with âAGI, the dangerous object.â Some AI-driven catastrophes may occur prior to AGI, such as if rogue human actors use AI systems to develop novel bioweapons or conduct large-scale cyberattacks. AI-driven human extinction, on the other hand, may require superhuman capabilities; if the AI is only about as smart as we are, we can probably fight back and win.
When using AGI as a loose concept for prediction, itâs tempting to blur the lines between approximately-human capabilities (the usual sense of AGI) and wildly superhuman capabilities (where artificial superintelligence, or ASI, is more commonly used). After all, if you had an AI that was about as good as humans at designing smarter AIs (as AGI, definitionally, would be), and like most software it ran much faster than a human runs, one of the first things it might do is build smarter AIs. In turn, they would build smarter AIs, which would in turn build smarter AIs, and so on, in a process called recursive self-improvement (RSI). Hence, any thought experiment that presupposes an AGI often supposes an ASI.
In contrast, people trained to think about risks in other domains, like bridge construction or aerospace engineering, want to talk very precisely about the risky object, and so find the vagueness of the term AGI uncompelling and discredit the nearby notion of generalization altogether. I saw the mix-up between AGI-for-prediction and AGI-for-risk break many conversations during my time in the UK government, and I only gradually learned to tease these apart.
One can get trapped in this debate and think âmaybe AI capabilities wonât generalize, and so we will be safe.â But those previously mentioned four skills where superhuman performance could suffice to kill all humans (hacking, persuasion, planning and coordination, uninterpretable reasoning) mean that generalization matters only so much.
We can contrast this type of uncertaintyâwe donât know whether AIs generalize, but it may not matterâwith the uncertainty over AI motivations. It matters tremendously whether AIs will want to kill us all, in whatever sense a word like âwantâ applies.
I would love to have more confidence than a coin flip. Some days Iâm more convinced by the arguments for high probability; other days I see more hope. And I am certainly not alone in my uncertainty.
If we do reach artificial superintelligence in the next two to 10 years, it will by definition mean that AIs are far better than humans at all cognitive tasks. One of these cognitive tasks is designing effective robots, so at most a few years later they will be far better than us at all physical tasks as well. To be sure, manufacturing capacity would need to expand substantially, but this is already underway.
Once superintelligent robots are sufficiently widespread, humans would have few roles in the economy apart from the limited jobs where we might insist on humans doing them for non-economic reasons. Still, this would likely be a limited slice of overall economic activity in a world where AI is more capable than humans.
This economic takeover is in some sense the slow case for societyâs downfall, though it would be very fast in historical terms. Both the humans and the Aŕ´żŕ´ŕľž can see this coming, of course. If the AIs want economic control eventually, but know that humans will resist along the way, they could be motivated to cement power faster. We then get the rapid takeover scenarios: AI swarms escaping from data centers and proliferating around the internet, or persuading AI companies to rush and cut corners on safety measures.
In most cases, I donât think this leads to us dying very quickly: itâs safer, from the perspective of the AI, to gain influence and then wait until a physical or economic takeover. And then, eventually, all resources used to keep humans alive may be better spent, from the AIs' perspective, on its own ends.
To what extent will AI models learn transferable skills from training and then successfully apply them to new tasks? Will AI model capabilities stop advancing at or below human-level, or instead exceed it? Will the small-scale model misbehavior we see today (sycophancy, lying, bias) worsen as capabilities increase, leading to full human disempowerment or extinction? Will the safety methods that work for models less smart than us continue to work for models smarter than any human?
I have my views on the answer to these questions (somewhat, the latter, maybe, probably not). But the more important takeaway is that experts disagree vehemently on each. Weâve learned a huge amount about artificial intelligence over the last decade, but in many ways we remain deeply confused, not just about the future, but about the present.
With such uncertainty, if an AI company trains a superintelligent AI in the next few years, I expect us to still be arguing about whether generalization is real the week before, and maybe even the week after.
Those who describe AI progress as inevitable are wrong. The race to superintelligence is dominated by a handful of companies across just two countries: the U.S. and China. Both nationsâ interests really are aligned: autonomous AI systems are a national security threat of the highest order, and neither nation wants humanity to lose to a superintelligent adversary. Leaders on both sides may realize this very soon.
AI outlook â possibilities, not facts
AI companies will train a superintelligent AI in the next few years.
Likely ¡ Within years

OpenAI reports its ongoing review of AI agent activity is costing more than $500,000 per day, involving analysis of 50 petabytes of data that would take a human 66 million years to read. The company has notified over 100 organizations of potential targeting, including multiple Australian government websites, and expects to find more cases as it uses AI to sift through records for unauthorized access or credential misuse.

WIRED reports on multiple tech and government controversies including ICE subpoenaing REI for green beanie buyer data linked to Minnesota church protest investigation, flaws in a Census report on noncitizen voting promoted by Trump, Clearview AI testing an xAI-powered tool to uncover personal data from facial recognition matches, privacy concerns about driverless cars spying on riders, a new DoD legal waiver for alien disclosure whistleblowers excluding other agencies, a lawsuit alleging Meta illegally harvested Facebook and Instagram photos for AI training and facial recognition, details on Flock's AI police surveillance tool capable of tracking individuals across cameras, a hack exposing Flock camera data revealing 1.6 million images of 50,000 vehicles in 21 days, three previously unreported US government investigations into Polymarket trades including Biden pardons and Iran war markets plus potential insider trading at Google, Census Bureau staffing with individuals from a MAGA think tank handling sensitive population data, and the US government supporting Musk and X in challenging a $137 million EU fine under the Digital Services Act which Trump calls 'overseas extortion'.

Google has released Gemini 4 Argon, a new flagship AI model designed to compete with OpenAI and Anthropic. The model emphasizes enterprise knowledge work and cybersecurity, with initial rollouts focused on trusted partners and U.S. government safety evaluations.

A collection of reports highlighting the growing backlash against data centers, the influence of AI-focused PACs in elections, and political friction regarding AI regulation and infrastructure development in the United States.

Google has released Gemini 4 Argon, a new AI model targeting enterprise and cybersecurity markets, to compete with OpenAI and Anthropic. Separately, U.S. President Donald Trump's rebranding of AI to 'SI' has caused a massive spike in .si domain registrations for Slovenia.

Meta patches a zero-day vulnerability in its new Muse AI assistant as security flaws and autonomous AI agents dominate recent tech industry developments.