What ARC-AGI measures, and what o3's high score does not prove
A benchmark of novel grid puzzles became the headline test for machine reasoning. OpenAI's o3 cleared the first version, but at thousands of dollars per task, and the harder ARC-AGI-2 shows how far models still sit from ordinary people.
By Zain
Published

A reasoning test built to resist memorization
The Abstraction and Reasoning Corpus, now called ARC-AGI, is a set of visual puzzles. Each task shows a few examples of a colored grid turning into another grid, then asks the solver to produce the output for a new input. The rule is never stated. You infer it from two or three demonstrations and apply it once.
Francois Chollet introduced the benchmark in his 2019 paper "On the Measure of Intelligence," where he defined intelligence as skill-acquisition efficiency: not how well a system performs a task it already knows, but how quickly it learns a new one from few examples (Chollet, 2019). ARC was built to reward that. The tasks rely only on "core knowledge" priors such as objects, counting, and basic geometry, and each puzzle is novel, so a model cannot win by having seen the answer in training. ARC Prize describes the goal as tasks "that humans solve effortlessly yet AI finds challenging" (ARC Prize).
That framing is the whole point. Most AI benchmarks measure crystallized skill. ARC-AGI tries to measure fluid reasoning on problems the model was not trained for.
o3 cleared the first version, at a price
For years the benchmark barely moved. ARC-AGI-1 went from 0% with GPT-3 in 2020 to about 5% with GPT-4o in 2024, according to ARC Prize. Then in December 2024, OpenAI's o3 reasoning model posted a result that made headlines.
On the semi-private evaluation set, o3 scored 75.7% in a high-efficiency configuration and 87.5% in a low-efficiency configuration that used roughly 172 times more compute (ARC Prize, 2024). On the larger public evaluation set the figures were 82.8% and 91.5%. For comparison, the earlier o1 model scored 32%. Chollet called the jump a genuine step up in capability while being blunt about its limits, which is where a lot of the coverage stopped reading.

The cost caveat nobody should skip
The score came with a bill. ARC Prize reported the high-efficiency semi-private result at about $26 per task and the 87.5% low-efficiency result at about $4,560 per task (ARC Prize, 2024). TechCrunch reported that the high-scoring configuration used more than $1,000 of compute per task, with the full test run costing far more (TechCrunch, 2024).
“A model can top ARC-AGI-1 and still fail puzzles a child solves, still hallucinate, and still need a disclaimer on its answers.”
That matters because efficiency is part of the definition being tested. A system that reaches human-level accuracy only by spending thousands of dollars and 172 times the compute per puzzle has not shown efficient general reasoning. It has shown that you can buy accuracy with search and compute. Compute prices do fall over time, and cheaper configurations have since reached comparable scores, but the 2024 headline number and its cost should always be read together.
ARC-AGI-2 reset the bar
Chollet's team saw this coming. Alongside the o3 result they said a second version, ARC-AGI-2, would launch in 2025 and would likely drop o3 under 30% while humans stayed above 95% (ARC Prize, 2024).
ARC-AGI-2 keeps the same input-output grid format but targets weaknesses the first version did not isolate: symbolic interpretation, compositional reasoning where several rules interact at once, and applying a rule differently depending on context (Chollet et al., 2025). It also keeps efficiency as an explicit axis.
The human calibration is the part worth holding onto. ARC Prize ran a live study in San Diego in early 2025 with more than 400 members of the public. Every task in the evaluation set was solved by at least two people in two attempts or fewer, and the average individual score was about 60% (ARC Prize; Chollet et al., 2025). These are not trick questions. Ordinary people solve them.
Machines, so far, do not. Through 2025 the top score on the ARC-AGI-2 private set in the public competition was about 24%, per ARC Prize's season review. The strongest verified commercial results were higher but still short of people: a refinement system built on Gemini 3 Pro reached 54% at about $31 per task, and Anthropic's Opus 4.5 in a thinking configuration reached 37.6% at about $2.20 per task (ARC Prize, 2025). The gap to a 60% human average, and to the 100% of tasks humans can solve between them, is the story.
What a score does and does not prove
Chollet has been consistent about the meaning of all this. "Passing ARC-AGI does not equate to achieving AGI," he wrote, "and, as a matter of fact, I don't think o3 is AGI yet. o3 still fails on some very easy tasks, indicating fundamental differences with human intelligence" (ARC Prize, 2024). He frames the benchmark as "a research tool designed to focus attention on the most challenging unsolved problems in AI," not a finish line.
So what does a high ARC-AGI score prove? It is real evidence that a system can generalize to some novel problems it was not trained on, which older models plainly could not. What it does not prove is that the system is generally intelligent, reliable, or efficient. A model can top ARC-AGI-1 and still fail puzzles a child solves, still hallucinate, and still need a disclaimer on its answers. And a benchmark is only a proxy: once a target is known, labs optimize for it, which is exactly why a harder version was needed within months.
The honest reading is narrower than the headlines. ARC-AGI is one of the better-designed tests of fluid reasoning, and progress on it since 2024 is genuine. It is not a thermometer for AGI, and the people who built it say so first.
Frequently asked questions
What does the ARC-AGI benchmark actually test?
It tests fluid reasoning on novel visual grid puzzles. You infer an unstated rule from a few examples and apply it to a new grid. Tasks use only core knowledge like objects and counting and are new each time, so a model cannot win by memorizing training data (ARC Prize; Chollet, 2019).
What did OpenAI's o3 score on ARC-AGI, and how much did it cost?
In December 2024, o3 scored 75.7% on the semi-private ARC-AGI-1 set at about $26 per task, and 87.5% in a configuration using roughly 172x more compute at about $4,560 per task (ARC Prize, 2024).
How do AI models do on the harder ARC-AGI-2?
Much worse. Average humans score about 60% and every task is solvable by people, but through 2025 the top public-competition score was about 24% and the best verified commercial result was 54% at around $31 per task (ARC Prize, 2025).
Does a high ARC-AGI score mean an AI has reached AGI?
No. Benchmark creator Francois Chollet says passing ARC-AGI does not equate to AGI and that o3 still fails some very easy tasks, showing fundamental differences from human intelligence. He calls it a research tool, not a finish line (ARC Prize, 2024).
Sources
- 1
OpenAI o3 Breakthrough High Score on ARC-AGI-Pub, ARC Prize (December 19, 2024)
- 2
ARC Prize 2025 Results and Analysis, ARC Prize (December 4, 2025)
- 3
ARC-AGI-2: A New Challenge for Frontier AI Reasoning Systems, arXiv (Chollet et al.) (May 16, 2025)