Superhuman intelligence: what it means, with examples from Go to chess
Superhuman describes a system that beats the best humans at a task. It does not mean superintelligence, which would exceed humans across nearly every domain. The difference is the one that matters for the race.
Published

What does superhuman intelligence mean? It means a system that performs a defined task better than the best human at that task. AlphaGo is superhuman at Go. A calculator is superhuman at long division. Neither is superintelligent, which is a much stronger claim: philosopher Nick Bostrom defines superintelligence as "any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest."
The two terms get mixed up constantly, and the mix-up shapes how people read AI news. A headline that says an AI is "superhuman" can describe a chess engine or a claimed step toward artificial superintelligence. This explainer sorts the cases, with dates, so you can tell which one you are reading.
The test: superhuman at what?
A useful definition has three parts: a task, a measure and a human baseline. Without all three, "superhuman" is marketing.
- Task: a bounded activity, such as playing Go, folding proteins or reading chest X-rays.
- Measure: how success is scored, such as win rate, accuracy or time to complete.
- Baseline: which humans, such as a median amateur or the best professional in the world.
A system can be superhuman against an average person and not against an expert, so the baseline matters as much as the score.
The famous example: Go
In March 2016 in Seoul, Google DeepMind's AlphaGo beat Lee Sedol 4 to 1, according to DeepMind. More than 200 million people watched. In the second game, AlphaGo played Move 37, which DeepMind says had a 1 in 10,000 chance of being used by a human player. Lee Sedol said afterward: "when I saw this move, I changed my mind. Surely, AlphaGo is creative." In the fourth game he answered with an equally unlikely Move 78 and won. AlphaGo was later awarded a 9 dan professional ranking, the highest available, the first computer system to receive it.
That is the textbook case. It was a rule-bound game with perfect information, a clear score and a human baseline of the best in the world. It was superhuman in the full sense, and it said nothing about whether the same system could, for example, run a lab.
“when I saw this move, I changed my mind. Surely, AlphaGo is creative.”

Superhuman without human data
DeepMind's AlphaZero pushed the point further. According to DeepMind's write-up of its December 2018 Science paper, AlphaZero trained for 9 hours in chess and beat Stockfish, winning 155 and losing 6 of 1,000 games. It searched about 60,000 positions per second, against Stockfish's 60 million. In Go it trained for 13 days and beat AlphaGo Zero in 61% of games.
Two features stand out. The system learned through self-play and reinforcement learning, not from hand-written rules, and it searched far fewer positions than traditional engines. Chess grandmaster Matthew Sadler described its style as pieces that "swarm around the opponent's king with purpose and power." The lesson is that superhuman performance can come from methods no human uses, and that it arrives fast once a training setup works.
Narrow versus general
Here is where the vocabulary matters. Wikipedia's summary of Bostrom's work separates narrow superhuman capability, already shown in domains such as chess or image recognition, from superintelligence proper. Bostrom also describes three forms of the latter: speed (human-level thinking running much faster), collective (many systems coordinating) and quality (reasoning that is fundamentally better). He notes that biological neurons peak at about 200 Hz, seven orders of magnitude slower than a roughly 2 GHz processor, which is one reason speed superintelligence is considered plausible.
Today's frontier systems sit in an awkward middle. They are superhuman at some tasks and below average at others, an uneven profile. Measuring how far that profile extends is the work of groups such as METR, whose time-horizon metric asks how long a task, as measured by human expert completion time, an agent can finish at a given reliability. METR says the trend is exponential over models from 2019 through 2025. It also cautions that its tasks are mostly software engineering, machine learning or cybersecurity, that they reflect what a low-context new hire could do, and that time horizons do not mean AI can automate jobs. Our guide to how AGI is measured covers those measures in detail.
What measures say about today's models
Measurement is where the vocabulary gets tested. METR's time horizon is the task duration, measured by how long human experts take, at which an agent is predicted to succeed at a given reliability. It is not the time the AI spends working. That makes it a difficulty measure, and METR says an exponential trend fits the data better than a linear or hyperbolic one. It is still a measure of a narrow band of work, and METR itself says it does not mean AI can automate jobs, because real work involves human interaction, subjective success criteria and messier conditions. So even a rising time horizon, taken alone, tells you a system is getting better at a family of tasks, not that it has crossed from superhuman at something to superhuman at everything.
“any intellect that greatly exceeds the cognitive performance of humans in virtually all domains of interest”
How to read a superhuman claim
When you see "superhuman" in a headline, run four checks.
1. Which task? If the answer is a game or a benchmark, ask whether the task resembles real work. 2. Which humans? Compare against experts, not averages. 3. Who measured it? A lab's own number is a claim; an independent evaluation is evidence. Our piece on what AI leaderboards rank explains why top scores can hide a lot. 4. Does it transfer? A system superhuman at Go cannot be assumed to be good at anything else.
How it connects to superintelligence
The jump from narrow to general is the central question of the race. Labs argue it is coming through scale and automated research, and skeptics argue the narrow wins do not add up to it. We keep the vocabulary straight in our AGI versus ASI glossary and our explainer on what superintelligence is.
Our position: use "superhuman" only with a task attached. The history from Go to chess shows machines passing the best humans quickly once the setup is right. It does not show that the setups generalize. Treat the first as settled and the second as the open question.
Frequently asked questions
What does superhuman intelligence mean?
It means performing a defined task better than the best humans at that task, as AlphaGo did at Go. It is narrower than superintelligence, which Bostrom defines as greatly exceeding human cognition in virtually all domains of interest.
Is AlphaGo superhuman?
Yes, at Go. Per Google DeepMind, it beat Lee Sedol 4 to 1 in March 2016 and received a 9 dan professional ranking. That performance is limited to Go and does not extend to other tasks.
What is the difference between superhuman intelligence and superintelligence?
Superhuman describes beating humans at specific tasks. Superintelligence describes an intellect that greatly exceeds humans across virtually all domains. Chess and Go engines are the former.
Are there AI systems that are already superhuman?
Yes, in narrow domains. AlphaZero beat Stockfish in chess after about 9 hours of training, per DeepMind's account of its December 2018 Science paper. Narrow superhuman performance has not been shown to generalize to all tasks.
How do researchers measure whether AI is getting more capable?
One method is METR's time horizon, the length of human expert task an agent can complete at a given reliability. METR says its tasks are mostly software, machine learning or cybersecurity and do not measure all intellectual work.
Sources
What each one is, and whose it is.
- 1
AlphaGo, Google DeepMind (March 14, 2016)
OtherThe vendor’s own - 2
AlphaZero: shedding new light on chess, shogi and Go, Google DeepMind (December 5, 2018)
Vendor announcement - 3
Superintelligence, Wikipedia (October 6, 2026)
OtherIndependent of the vendor - 4
Task-completion time horizons of frontier AI models, METR (May 7, 2026)
BenchmarkIndependent of the vendor