What the AI leaderboards rank in late 2026, and what a top score hides

Frontier model leaderboards in late 2026 crown a single winner, but the rank is a weighted average that hides contamination, task-specific weakness, and large cost and speed gaps. Here is what a top score does not tell you.

By Yash Malviya

Published

A dimly lit desk setup displaying cryptocurrency trading charts on multiple screens
Photo: Rafael Minguet Delgado / Pexels

Ask which AI model is best in late 2026 and most people point to a leaderboard. Artificial Analysis publishes a single Intelligence Index. LMArena ranks models by human votes. Benchmark aggregators stack them by percentage scores. Each produces a tidy ordering, and the model on top gets treated as the winner. The ranking is real. The problem is what it leaves out.

Who leads, and by how much

As of early October 2026, Artificial Analysis put Claude Opus 5.5 at the top of its Intelligence Index at roughly 58, with OpenAI's GPT-6 Astra and Google's Gemini 4 Argon close behind, according to the company's public leaderboard. The top of the board is entirely closed. The strongest open-weight models, led by Chinese labs including Xiaomi's MiMo, Z.ai's GLM-5.3, and Moonshot's Kimi K3, scored about 44 to 46 on the same index. That gap, closed models near 58 against open models in the mid-40s, is one of the few things a single number conveys honestly: the frontier is still proprietary, and the distance is not small.

A score is an average of someone's choices

A headline index number is not a measurement of one thing. It is a weighted blend. Artificial Analysis says its Intelligence Index v4.3.2 combines ten evaluations across four groups: agentic tasks at 30 percent, coding at 20 percent, general knowledge and reasoning at 30 percent, and scientific reasoning at 20 percent. Change those weights and you change the winner.

The individual benchmarks test narrow skills. MMLU-Pro is a multiple-choice knowledge test with ten options per question. GPQA Diamond is a set of a few hundred PhD-level science questions. Humanity's Last Exam, released in January 2025 by the Center for AI Safety and Scale AI, collects a couple thousand hard academic questions across many fields. LiveCodeBench pulls fresh programming-contest problems over time. SWE-bench Verified asks a model to fix 500 real GitHub issues. A model can top the overall index and still sit mid-pack on the one benchmark that matches your actual work.

“A leaderboard rank is an average of other people's weighting choices. It is not a fact about your task.”

Close-up of a magnifying glass over financial data charts and metrics on printed paper
An index score collapses many different tests into one number, which can hide exactly where a model is weak. Photo: RDNE Stock project / Pexels

When the test leaks into the training data

The deeper problem is that a high score can measure memorization. Benchmarks are public, and model training data is scraped from the public web, so the test often ends up inside the training set. Meta's own Llama 2 report in 2023 found that more than 16 percent of MMLU test items overlapped with its training data, with about 11 percent seriously contaminated.

The effect shows up when someone builds a clean test. Scale AI wrote a new grade-school math set, GSM1k, to mirror the popular GSM8k, then compared scores. Some model families, notably Phi and Mistral, fell by as much as 13 points on the fresh questions, while frontier models from OpenAI and Anthropic barely moved, Scale reported in 2024. Even SWE-bench Verified, the coding benchmark everyone cites, has faced questions about whether its Python tasks appeared in training data. LiveCodeBench tries to dodge this by only counting problems published after a model's training cutoff, which is the right instinct but not a complete fix.

Benchmarks that break the moment they matter

Benchmarks also wear out. MMLU was meant to be a durable expert-knowledge test when it launched in 2020, when GPT-3 scored about 43 percent. By 2023, OpenAI reported GPT-4 at 86.4 percent, near expert level. Once a benchmark saturates, it stops telling you who is better, because everyone scores near the ceiling.

This is Goodhart's law in action: when a measure becomes a target, it stops being a good measure. Labs optimize for the tests the field watches, the tests saturate, and the field moves to harder ones. The churn is visible inside the Intelligence Index itself. Earlier versions leaned on MMLU-Pro, GPQA, and math contests; the late-2026 version has shifted weight toward agentic and long-horizon tasks as the older tests lost their power to separate models.

The leaderboard can be played

Human-preference rankings have a different weakness. A 2025 study titled The Leaderboard Illusion, from researchers at Cohere Labs, Stanford, and Princeton, examined roughly two million Chatbot Arena battles and argued the game is not evenly played. They found large labs could test many private variants and publish only the best one; they counted 27 private variants Meta tried before the Llama 4 release. They estimated Google and OpenAI each received around 19 to 20 percent of all arena data, while 83 open-weight models together got under 30 percent. Access to that data, they calculated, could lift a model's arena win rate substantially. The arena also rewards style: longer answers and bulleted lists tend to win votes whether or not they are more correct.

What to do with a ranking

None of this makes leaderboards useless. They are a fast way to see roughly where a model sits and to confirm, for instance, that open weights still trail the closed frontier. They are a bad way to choose a model for a job. The single number hides cost and speed, which vary enormously: Artificial Analysis lists per-task costs from about one cent to several dollars across model variants, and output speeds that differ by an order of magnitude. A rank also hides where a model is weak, since it is an average.

The practical read is simple. Use the top of a leaderboard as a shortlist, not a verdict. Then look at the sub-scores for the task you care about, check whether the benchmark could be contaminated, and price out the latency and cost before you commit. A leaderboard rank is an average of other people's weighting choices. It is not a fact about your task.

Frequently asked questions

What do AI model leaderboards actually measure?

They measure a weighted blend of narrow tests, not one thing. The Artificial Analysis Intelligence Index v4.3.2 combines 10 evaluations across agents, coding, general reasoning, and science, each with a fixed weight. LMArena instead ranks models by human votes on pairs of answers. Change the weighting or the voting pool and the order can change.

What is benchmark contamination?

It is when test questions end up in a model's training data, so a high score can reflect memorization rather than skill. Meta's Llama 2 report found over 16% of MMLU items overlapped with training data (2023), and Scale AI's clean GSM1k test showed some models dropping up to 13 points versus the public GSM8k set (2024).

Which AI model is best in 2026?

There is no single best. On the Artificial Analysis Intelligence Index in early October 2026, closed models led, with Claude Opus 5.5 near 58 and OpenAI and Google close behind, while the strongest open-weight models scored about 44 to 46. The right choice depends on your task, your cost ceiling, and your latency needs.

Why do AI benchmarks saturate so quickly?

Because labs optimize for the tests the field watches. MMLU launched in 2020 with GPT-3 near 43%; by 2023 OpenAI reported GPT-4 at 86.4%, close to the ceiling. Once scores bunch near the top, the benchmark can no longer separate models, an example of Goodhart's law, and the field moves to harder tests.

Sources

  1. 1

    The Leaderboard Illusion, Cohere Labs, Stanford University, Princeton University (arXiv) (April 28, 2025)