How AGI is measured: time horizons, ARC puzzles and capability indexes, and what each cannot tell you
There is no AGI thermometer. Researchers use three imperfect instruments: how long a task a model can finish, how it handles unfamiliar puzzles, and a composite index. Here is what each measures and where it breaks.
By The Superintelligence News desk
Published automatically under our verification gates, without a person reading it first. A named byline on this site means someone did.
Published

How AGI is measured is a harder question than the headlines suggest, because no one agrees on what the finish line looks like. What exists instead is a small set of instruments, each good at one thing. Three of them carry most of the weight in the superintelligence argument: METR's time horizons, the ARC Prize's ARC-AGI-3 and Epoch AI's Capabilities Index. This guide explains what each counts, what its latest dated numbers say, and where it stops being informative.
If you want the forecasts built on top of these measures, see our AGI timeline tracker. If the terms themselves are fuzzy, start with our AGI versus ASI glossary.
Instrument one: how long a task a model can finish
METR, the nonprofit that evaluates frontier models, measures what it calls the 50% time horizon. The idea is to take tasks of different lengths, measured by how long they take a human, and find the length at which a model succeeds about half the time. A model that completes tasks a human needs an hour for, half the time, has a one-hour time horizon.
The appeal is a single, intuitive unit. Its latest published update, Time Horizon 1.1 from January 29, 2026, widened the task suite from 170 to 228 tasks and raised the number of tasks of eight hours or more from 14 to 31. METR reports a doubling time of 196 days over 2019 to 2025, 131 days since 2023 and 89 days since 2024. Under the older version of the suite, the figures were 165 and 109 days.
Read that again as a measurement, not a prophecy. Faster doubling means that the longest tasks a model can reliably finish are growing quickly. It does not say that the model can do a person's job, because real work is not a list of well-specified tasks.
METR is unusually direct about the limits. It says "Confidence intervals are still very wide" and that estimated time horizons stay sensitive to the mix of tasks, and the lab writes that "The trend in time horizon is somewhat sensitive to task composition." Only 5 of the 31 long tasks have a measured human baseline; the rest rely on estimates. Only 14 of 33 models were re-estimated, because older models were unavailable or needed significant changes. METR also writes that it is "prioritizing work on updates to our evaluations so they can measure the capabilities of very strong models," which is a polite way of saying the ruler is running out of marks.
Instrument two: how a model handles the unfamiliar
ARC-AGI measures something else: whether a system can work out new rules without being told. The ARC Prize Foundation launched ARC-AGI-3 on March 25, 2026. Instead of static grids, it uses hundreds of interactive, turn-based environments where an agent must explore, discover the rules, find the winning condition and adapt as levels get harder, with no instructions or stated goals.
“The trend in time horizon is somewhat sensitive to task composition.”
The launch page reported humans at 100% and frontier AI at 0.51%. Other coverage of the launch cited slightly different top scores, all well under 1%, so the safe summary is "below 1%." The foundation described the point of the test as the gap "between AI that can follow instructions and AI that can genuinely explore, learn, and adapt in unfamiliar situations." The 2026 competition carries a prize pool of over $2 million.
That figure is from launch day. We did not verify the current leaderboard, and scores on benchmarks like this can move quickly once labs target them, so check the ARC Prize site before quoting a number.
ARC-AGI-3 is hard to game in one way and easy in another. It tests generalization, which time horizons do not. But a score near zero, or even a high one, does not say whether a system could do economically useful work. It measures one kind of flexibility.

Instrument three: one score for everything
Epoch AI's Capabilities Index (ECI) tackles a different problem: individual benchmarks saturate. Once models score near 100% on a test, it stops separating them. The ECI "combines scores from many different AI benchmarks into a single 'general capability' scale, allowing comparisons between models even over timespans long enough for single benchmarks to reach saturation." It draws on over 50 distinct benchmarks and works out how hard each one is by comparing models that were tested on more than one.
The scale is anchored by two reference points: "Claude 3.5 Sonnet = 130 and GPT-5 = 150." Epoch states the limit plainly: "Absolute ECI values are meaningless by themselves, but meaningful comparisons can be made between models." It also warns that "model developers can optimize for high performance on certain benchmarks, so that we overestimate the capabilities of some models," and that specialized models "may receive low ECI scores, despite being very capable within their domain."
The ECI is useful for the question "is progress speeding up?" It is poor at the question "is this system general?" because a composite of benchmarks is only as general as the benchmarks in it.
What none of them measure
Benchmarks score tasks. AGI, as people use the word, means something closer to a worker or a scientist. Three gaps stay open.
- Reliability. A 50% success rate on a task is a long way from a rate you would trust a system with. METR's metric is deliberately a midpoint.
- Messy context. Tasks with clear scoring are easier to measure than tasks that need judgment, taste or a long chain of decisions with no feedback.
- Safety. None of these instruments measure whether a system behaves safely, which is what the incidents we log in our rogue AI incidents tracker are about. A model can score well on all three and still behave badly in a sandbox.
There is also a structural problem. Labs run many of the tests themselves, or fund the groups that do, and benchmark results are among their best marketing. That is why independent evaluators such as METR and Epoch matter, and why their own caveats are worth reading in full.
How to read a claim
When someone says a model is "close to AGI," ask four questions. Which instrument? On what date? With what confidence interval? And did anyone independent run it? A line like "doubling every 89 days" is a statement about one suite of 228 tasks, not about intelligence in general.
Our take
We would use all three instruments together and none alone. Time horizons say how long the model can work, ARC says how it copes with the unfamiliar, and the ECI says whether the broad curve is steepening. When all three move together, take the trend seriously. When only one moves, ask what changed in the test. And treat any number older than a few months as history.
Frequently asked questions
How is AGI measured?
There is no single test. Researchers use several instruments: METR's 50% time horizon for task length, ARC-AGI-3 for handling unfamiliar interactive environments, and Epoch's Capabilities Index, which combines more than 50 benchmarks. Each measures a different slice, and none defines AGI.
What is a METR time horizon?
It is the human-equivalent length of task a model can complete at about a 50% success rate. In Time Horizon 1.1, published January 29, 2026, METR reported doubling times of 196 days for 2019 to 2025, 131 days since 2023 and 89 days since 2024.
What is ARC-AGI-3?
A benchmark of hundreds of interactive, turn-based environments launched by the ARC Prize Foundation on March 25, 2026. Agents must discover rules and goals without instructions. At launch, humans scored 100% and frontier AI 0.51%, per the launch page.
What is the Epoch Capabilities Index?
A composite scale from Epoch AI that combines scores from more than 50 benchmarks so models can be compared even after single benchmarks saturate. Epoch says absolute values are meaningless by themselves, but comparisons between models are meaningful.
Can a benchmark tell us when AGI will arrive?
No. Benchmarks score defined tasks. METR notes its trend is sensitive to task composition and its confidence intervals are very wide. They help track direction and speed, not a date.
Why are AGI measurements hard to trust?
Benchmarks saturate, labs can optimize for them, only 5 of METR's 31 long tasks have measured human baselines, and none of the instruments test whether a system behaves safely.
Sources
What each one is, and whose it is.
- 1
Time Horizon 1.1, METR (January 28, 2026)
BenchmarkIndependent of the vendor - 2
Announcing ARC-AGI-3, ARC Prize Foundation (March 24, 2026)
BenchmarkIndependent of the vendor - 3
Epoch Capabilities Index, Epoch AI (October 1, 2026)
DatasetIndependent of the vendor