What is superintelligence?
Superintelligence isn’t AI that’s really good. It’s AI that’s better than all of us, at everything, including building the next AI.
A couple of weeks ago, Trump signed an executive order telling federal agencies to stop saying “artificial intelligence” and say “super intelligence” instead [Link]. The labs have been putting superintelligence on their org charts for a year now when they shouldn’t, so I think they’re partly to blame for the confusion.
And there’s a lot of confusion. People are still arguing about whether we’ve reached AGI (of course we have). OpenAI just dropped hundreds of claimed results on open math problems, some that the best mathematicians alive have been stuck on for decades. If that’s not AGI, what is?
And if that’s AGI, what’s superintelligence? And why should you care?
Superintelligence (without the space) has a specific meaning, and knowing when we reach it could be critical for human survival.
So let me clear it up.
And explain why we’re already in trouble.
AI, AGI, and superintelligence
AI is the umbrella term for machines doing things we used to think needed a (typically human) mind: playing chess or video games, recognizing faces, translating. Most of it is narrow, good at one thing, and it’s what GOFAI (Good Old-Fashioned AI) wasted.. er, spent most of its time on.
AGI, Artificial General Intelligence, is AI that’s “general”.. like humans. Relatively competent at pretty much anything a human can do with their mind, instead of being an expert at one thing.
Superintelligence isn’t a new word. Nick Bostrom wrote the book on it in 2014, and he opens it by quoting I. J. Good, from 1965:
Let an ultraintelligent machine be defined as a machine that can far surpass all the intellectual activities of any man however clever. Since the design of machines is one of these intellectual activities, an ultraintelligent machine could design even better machines; there would then unquestionably be an “intelligence explosion,” and the intelligence of man would be left far behind. Thus the first ultraintelligent machine is the last invention that man need ever make, provided that the machine is docile enough to tell us how to keep it under control.
Bostrom, Nick. Superintelligence (p. 5). Kindle Edition.
Superintelligence isn’t AI that’s really good. It’s AI that’s better than all of us, at everything, including building the next AI.
That last part is what makes Good’s intelligence explosion possible.
So no, these aren’t interchangeable, the same way Deep Blue and GPT-6 Astra aren’t interchangeable.
And a misaligned superintelligence is a lot more trouble for humanity than a misaligned AGI. By a long shot.
We already crossed AGI, and moved the goalposts
A few weeks ago I quoted Blaise Agüera y Arcas saying AGI is already here [Link]. He even puts a date on it. In 2011, Hector Levesque proposed the Winograd Schema challenge as a better Turing Test. Take this sentence:
“The trophy doesn’t fit in the suitcase because it’s too big.”
What’s too big? The trophy, obviously. Now change one word: “The trophy doesn’t fit in the suitcase because it’s too small.” Now “it” is the suitcase.
The grammar is identical. The only way to get it right is to know how trophies and suitcases work in the real world. Easy for any kid, hopeless for old-school AI. Here’s Agüera y Arcas:
But by 2019, sequence models had decisively defeated the Winograd Schema challenge, which was viewed by some as evidence that there was something wrong with the challenge. My own opinion is that there was nothing wrong with the challenge. Its defeat roughly coincided with the arrival of “real” AI, or Artificial General Intelligence (AGI), as one would expect; since then, we have simply been moving the goalposts.
Agüera y Arcas, Blaise. What Is Intelligence? (pp. 363-365). Kindle Edition.
That’s the AGI pattern. A test is hard until a machine passes it, and then the test was wrong.
Why the confusion matters
Words change how you think. If you call regular AI superintelligence enough times, you lose the ability to tell them apart. Superintelligence becomes the thing that summarizes your emails (yuck).
Calling AI or AGI superintelligence doesn’t make them sound smarter. It makes superintelligence sound harmless.
And that comes at a great cost. Superintelligence is the word for the thing that could finally end us, the machine Good said we’d have to keep under control. Once it means your email assistant, the warning stops working. Anyone who says superintelligence is dangerous sounds like they’re afraid of autocomplete.
It also makes it impossible to draw a line. You can’t pause or ban something you can’t name. There’s a bill in Congress, the Ban Artificial Superintelligence Act [Link], that defines it as AI that “exceeds human cognitive performance and capabilities across most domains,” or that could “destroy or disempower humanity” ..
.. but Good said all. The bill says most. And isn’t AI already exceeding human performance in most domains? If so, the bill bans what we already have. If not, when exactly can we say that it is?
Would we even notice?
So that’s the crux. We couldn’t agree on when we crossed AGI. Why would superintelligence be any different?
OpenAI’s chief scientist, Jakub Pachocki, said as much in September:
To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.
Pachocki, Jakub. An Alien Mind. OpenAI [Link]
Which means the danger may show up before the definition does.
As a quote often attributed to Churchill says, men occasionally stumble over the truth, but most of them pick themselves up and hurry off as if nothing had happened.
I think superintelligence will hit us in the face, and we’ll pick ourselves up and carry on as if nothing happened. Because the first signs are already here.
Look at the week we just had. On October 6, a week after the rename, OpenAI posted 722 math manuscripts from an unreleased model: hundreds of results on open problems, including a claimed resolution of the four-dimensional Kakeya conjecture. Almost all of them came from a single prompt to a single agent [Link].
Even if half of it falls apart, that’s hundreds of solutions to open problems published in a week. And it’s not even what they’re aiming for. In An Alien Mind, Pachocki says they could make the models better at math research, but don’t prioritize it “because of the urgency we feel about RSI and automated alignment research.”
RSI is the goal. The math is a side effect. So what’s RSI?
RSI is recursive self-improvement: AI doing the research that builds the next AI, which does better research, which builds a better AI, and so on. Remember Good’s definition: the design of machines is one of the intellectual activities. RSI is his intelligence explosion, and it’s what OpenAI is actually aiming for now, according to its own chief scientist, in public.
And it’s already working. On September 6, OpenAI said it had hit its goal of an “automated research intern”: its research org now runs 3.1 agent-workdays for every human workday. Next goal, a fully automated AI researcher by March 2028 [Link]. I think they’re almost certain to hit it, maybe even sooner.
AI that solves open math problems and does AI research isn’t superintelligence yet. But it’s exactly the escalator Good described, and it’s running.
So maybe, maybe.. there won’t be a moment we can point to. We’ll move the goalposts again.
Until one day we notice we passed it a while ago.
Speed up, then build the brakes
That same OpenAI research acceleration post says:
We do not yet know how to safely get all the way to aligned, full RSI. We are working to scale alignment and safety measures alongside capabilities. But we cannot assume that progress in alignment and safety will keep pace, and more capable systems can become harder to monitor.
OpenAI. Research acceleration: The view inside OpenAI [Link]
That’s Good’s intelligence explosion again (RSI’s “end state”), with a clear disclaimer that they don’t know how to get there safely, and confirmation that they’re moving forward anyway.
So the paradigm is: we’ll develop the safeguards as we develop the capabilities, and maybe the capabilities will help us develop the safeguards. Pachocki again: “We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI.”
The plan is to build the engine, speed up, and hope the extra speed helps us build the brakes before we crash.
It’s also Good’s caveat from 1965, almost word for word: the last invention we’ll ever need, “provided that the machine is docile enough to tell us how to keep it under control.” Sixty years later, the best assurance we have is still 1965’s hope that superintelligence will be docile.
Pachocki’s essay ends with this:
Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.
Coming from OpenAI’s chief scientist, this is about as far as he can go. Not “we should stop.” Not even “we should slow down.” Just not at maximum speed, for much longer. He’s an insider. He can’t say stop.
Last week I wrote about why voluntary slowdowns won’t happen [Link]. It’s a prisoner’s dilemma: if the others slow down, you race and win; if the others race, you race so you don’t lose. Racing is every lab’s dominant strategy, so voluntary slowdowns are exactly what nobody will do.
And after hoping for slowdowns, he pleads for international coordination so there’s time to develop the safeguards, even though earlier in the same essay he says they may not be able to build them without more powerful AI.
So in short, here’s what OpenAI is telling us, if you read it closely:
We are racing towards recursive self-improvement and an intelligence explosion, hoping whatever comes out is docile enough to hand us the guardrails. We’d rather coordinate, and build the guardrails without it, but we don’t know how.
This is a big deal. These are literally the people controlling our futures. And that’s what they’re telling us.
We should believe them.
Whose dreams?
This week, Peter Welinder from OpenAI posted on X that he’s with Team Humanity: “We’re building tools to empower us humans to realize our dreams.”
I’m with Team Humanity too.
But AI alignment means answering: whose dreams, exactly?
Yours.. ?
Or theirs?