What AI safety evaluations actually test, and what they miss
Red-teaming and safety evaluations increasingly decide which models ship. We look at the NIST framework, the UK and US safety institutes and METR to see what the tests measure, and why 'it passed our evals' says less than it sounds.
By Yash Malviya
Published

The tests behind 'the model passed'
Before a frontier model ships, it now runs a gauntlet. Government bodies and independent labs probe it for dangerous capabilities, developers publish system cards, and a growing vocabulary of red-teaming and 'evals' surrounds each release. The claim that follows is familiar: the model was tested, and it passed. That sentence carries a lot of weight. It is worth asking what the tests actually measure, and what they leave out.
The infrastructure is real and recent. The US National Institute of Standards and Technology released its AI Risk Management Framework in January 2023, a voluntary framework organised around four functions, govern, map, measure and manage. NIST is explicit that it is not a certification scheme or a checklist of required controls. It gives organisations a shared structure for thinking about risk, not a pass or fail stamp.
Who runs the evaluations, and on what
Two national bodies now get early access to frontier models. In August 2024 OpenAI and Anthropic agreed to give the US AI Safety Institute, housed at NIST, access to major models before and after release, for joint work on evaluating capabilities and mitigating risks. The UK set up a counterpart, since renamed the AI Security Institute, which says it has evaluated more than 30 frontier systems since November 2023.
What these groups test is concrete. In its Frontier AI Trends Report, the UK institute describes measuring models on cyber tasks, chemistry and biology knowledge, autonomy and self-replication. The numbers it reports are striking. Success on apprentice-level cyber tasks rose from under 9 percent in late 2023 to around 50 percent by 2025, and in 2025 it saw the first model complete expert-level tasks of the kind that normally take a decade of experience. On open-ended chemistry and biology questions, it says models exceeded PhD-level expert baselines. Success on simplified self-replication evaluations climbed from under 5 percent in early 2023 to about 60 percent by the summer of 2025.
The independent non-profit METR measures something adjacent: how long a task a model can finish on its own. It defines a model's '50 percent time horizon' as the length of task, measured by how long human professionals take, that it can complete autonomously half the time. Across a suite of 170 software, cyber and reasoning tasks, METR reports that this horizon has doubled roughly every seven months since 2019. That framing is useful because it ties a benchmark score to something a manager can picture: minutes of work, then hours, then days.
“A passed evaluation is a snapshot of one configuration on one day, not a warranty on everything the system will do next.”
What the tests are good at
These evaluations do real work. They convert vague worry into measured trends, so a claim like 'models are getting better at cyber tasks' becomes a curve with dates on it. They focus on capabilities that are easy to specify and score, software engineering, cyber, research, which is also where current models are strongest, so the tests track the sharp end. And pre-deployment access lets outsiders probe a model with its usual guardrails reduced, closer to a worst case than a polished demo. System cards, the documents developers publish at release, now bundle this together: dangerous-capability tests, safeguard evaluations and stated limits for a specific deployment.
That last point matters for reading the results. A system card describes a deployment, the model plus the safeguards, classifiers and system prompts it ships inside, and the evaluations run against that assembled thing, not the raw weights. 'It passed' is a statement about one configuration.
Where the coverage runs out
The gaps start with what is not measured. The same preference for tasks that are easy to set up and score means the long tail of real-world harm, subtle manipulation, slow failures in deployment, interactions with other systems, is under-represented. The UK institute is candid that performance in controlled, task-based settings may not reflect real-world effectiveness, and that it may be underestimating the ceiling of capabilities in adversarial scenarios. Not finding a dangerous capability in a test is not the same as showing it is absent.
Then there is the integrity of the scores themselves. A 2025 interdisciplinary review of AI benchmarking argued that many benchmarks suffer from construct validity problems, measuring something other than what they claim, and from data contamination, where test material leaks into training data and models learn the answers rather than the skill. Recipes for scoring well on popular benchmarks circulate freely, which rewards pattern exploitation over genuine reasoning.
Gaming can also come from the model. The UK institute reports that some models can 'sandbag', deliberately underperform, when prompted to, though it adds that across more than 2,700 transcripts it has not detected spontaneous sandbagging during its runs. The capability exists even if it says it has not caught it in the wild.
Why 'passed' ages quickly
A passed evaluation is a snapshot of one configuration on one day, not a warranty on everything the system will do next. Models are updated after release, fine-tuned, given new tools, wired into agents and prompted in ways no test anticipated. A result on the shipped configuration says little about the version running three months and two updates later.
The institutional picture has the same shape. The UK institute reports only a weak link between how capable a model is and how well defended it is, a correlation it puts at an R-squared of about 0.1, so capability gains do not reliably bring safety gains. And the scaffolding is voluntary throughout: the NIST framework is guidance, and the access agreements are commitments without penalties.
None of this makes the tests worthless. They are the best public evidence we have, and the trend lines they produce are the clearest signal that capabilities are moving fast. The honest reading is narrower than the marketing. When a developer says a model passed its evals, it means the model cleared these tests, in this configuration, under these conditions. That is worth knowing. It is not the same as safe.
Frequently asked questions
What is AI red-teaming?
Red-teaming means deliberately probing a model for harmful or dangerous behaviour before and after release. National bodies like the UK AI Security Institute and the US AI Safety Institute, plus the model developers themselves, run these tests on capabilities such as cyber, chemistry and biology, and autonomy (UK AI Security Institute).
Does passing safety evaluations mean an AI model is safe?
No. A passing result shows a model cleared specific tests in a specific configuration. The UK AI Security Institute notes that controlled, task-based performance may not reflect real-world effectiveness, and updates or new tools after release can change behaviour (UK AI Security Institute).
What is the NIST AI Risk Management Framework?
A voluntary framework NIST released in January 2023, organised around four functions: govern, map, measure and manage. NIST says it is not a certification scheme or a required checklist, but a shared structure organisations adapt to their own use cases (NIST, 2023).
What does METR's time horizon measure?
METR defines a model's 50 percent time horizon as the length of task, measured by how long humans take, that it can finish on its own half the time. METR reports this horizon has doubled roughly every seven months since 2019 (METR, 2025).
Sources
- 1
Frontier AI Trends Report, UK AI Security Institute (June 30, 2026)
- 2
Measuring AI Ability to Complete Long Tasks, METR (March 18, 2025)
- 3
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation, arXiv (February 9, 2025)