UK AI Security Institute vs CAISI: what each tests, and what powers they have

The UK AI Security Institute and the US CAISI both test frontier models, but their budgets, mandates and published work differ sharply. Neither can block a release.

By Yash Malviya

Published

A frontal view of the iconic US Capitol Building in Washington D.C. under blue skies
Photo: Guohua Song / Pexels

The UK AI Security Institute and the US Center for AI Standards and Innovation (CAISI) are the two governments' main testers of frontier models, and they are often treated as twins. They are not. One is a research organization with a ministerial budget line and a stated pipeline into pre-release models. The other is a small office inside a standards agency that lives or dies on voluntary agreements and annual appropriations. This explainer sets out what the UK AI Security Institute versus CAISI comparison actually looks like on October 5, 2026, using each body's own pages.

The short version

  • UK AISI sits inside the Department for Science, Innovation and Technology. It describes its job as informing governments so they can keep the public safe, and it lists pre-deployment access to leading models among its resources.
  • CAISI sits inside NIST, the US standards agency. Its page describes it as industry's primary point of contact within the US government for testing and collaborative research on commercial systems.
  • Neither is a regulator. Both rely on cooperation from developers. Neither page claims a power to block a release.

What each one says it tests

AISI's about page lists its research areas as cyber misuse, safeguards, alignment, control, autonomy, human influence and societal resilience. The autonomy line is blunt: it asks how capable systems are at conducting their own AI research, making copies of themselves and evading attempts to control them. That is a long way from a benchmark leaderboard.

CAISI's mandate, as its NIST page words it, is narrower and more national-security shaped. It is to lead unclassified evaluations of capabilities that may pose risks to national security, focusing on "demonstrable risks, such as cybersecurity, biosecurity, and chemical weapons." It also has a second job AISI does not advertise: assessing adversary systems, including the possibility of backdoors and covert behavior, and tracking international competition.

One naming note. The page header at nist.gov/caisi currently reads "Center for Advancing Innovation and Standards for Super Intelligence (CAISSI)" while NIST's own news posts from 2026 still say Center for AI Standards and Innovation. We use CAISI, the name in the dated evaluation posts, and flag that NIST's naming is in flux.

Two researchers in a laboratory examining data on a computer and tablet, wearing safety gear
Government researchers reviewing evaluation results at a workstation. Photo: https://kaboompics.com/ / Pexels

Powers: persuasion, not subpoena

AISI says it tests leading systems before they are released and collaborates with top companies on safety. That wording describes access by arrangement. CAISI is more explicit: it says it will establish "voluntary agreements" with private developers and evaluators. Voluntary is the operative word on both sides of the Atlantic. If a lab declines, neither body's published mandate gives it a stated way to compel access.

This is why the question of who tests the testers matters. Our earlier piece on the GPT-6 Astra supply-chain attack findings covers a case where AISI's own evaluation setup was the story.

“I believe that the amount of resources truly needed when the dust settles here will be closer to $100 million a year”

Rep. Jay Obernolte, via FedScoop, July 23, 2026

Staff and money, as far as it is public

The gap is large, and it is partly a gap in disclosure.

AISI's about page claims 100+ technical staff, funding of 66 million pounds per financial year, priority access to over 1.5 billion pounds of compute through UK research programs, and the ability to mobilize more than 15 million pounds in grants. Those are the institute's own figures, undated on the page, so treat them as a snapshot.

CAISI has no comparable page. What is public is the money. Federal News Network reported on January 5, 2026 that Congress funded NIST's AI research and measurement work at $55 million, with up to $10 million for CAISI's expansion. FedScoop reported on July 23, 2026 that CAISI's fiscal 2026 funding was $10 million, the same as the prior two years, and that the current House bill provides $20 million. We found no published CAISI headcount, so we do not give one.

At the hearing FedScoop covered, Representative Jay Obernolte said the amount truly needed "will be closer to $100 million a year." OSTP Director Michael Kratsios said CAISI "can certainly benefit from more resources." Even without a dated currency conversion, AISI's 66 million pounds is plainly larger than CAISI's $10 million, though the two budgets cover different scopes.

What each has published recently

CAISI publishes its evaluations as short NIST news posts, and recent ones focus on Chinese open-weight models:

  • May 1, 2026: DeepSeek V4 Pro. CAISI said its capabilities lag the frontier by about 8 months, and that the model scored better on DeepSeek's self-reported evaluations than on CAISI's, which include non-public benchmarks.
  • July 23, 2026: a joint UK AISI and CAISI preliminary assessment of Kimi K3's cyber capabilities. In a simulated corporate network attack, Kimi K3 reached step 17 of 32 on average, against 28.5 for the most cyber-capable US models. The post also says its safeguards did not stop exploit development attempts.
  • September 17, 2026: Z.ai's GLM-5.3. CAISI called it the most cyber-capable open-weight model released to date, and about four months behind the US frontier on its aggregate cyber measure.

The Kimi post carries a detail worth noticing: US closed-weight models were tested with system-level safeguards disabled, to measure maximal capability. Public versions have safeguards on.

AISI's output reads differently. Its research page lists dozens of papers and blogs by theme, from alignment and control to red teaming. The latest, dated September 28, 2026, is "Evaluating Whether GPT-6 Astra Performs Unsanctioned Supply-Chain Attacks." Earlier entries include an April 27 study of whether models would sabotage safety research. It also maintains the open-source Inspect evaluation tool and published a Frontier AI Trends Report in December 2025.

Where they overlap

The Kimi K3 post proves the two can work as a pair. The pattern otherwise splits cleanly. CAISI's public record is mostly benchmark scorecards on foreign models, framed around US competitiveness. AISI's is a broader research program that includes behaviors like sabotage and deception. If you want to know how far behind China is on cyber, read CAISI. If you want to know whether a frontier model will quietly work against its overseers, AISI is publishing more.

Our take

Treat them as complements, but do not mistake either for oversight with teeth. AISI has the budget, the stated pre-deployment access and a research agenda that reaches the hard questions. CAISI has the standards agency's convening role, a national-security mandate and a thin purse that Congress is only now debating doubling. The honest gap is power: both can measure, neither can say no.

What we would watch is the fiscal 2027 appropriation for CAISI, whether the voluntary agreements it promises are ever published, and whether CAISI's rename sticks. We would also like CAISI to publish a headcount and AISI to date its numbers. A tester that will not say how big it is invites the same scrutiny it applies to labs.

Frequently asked questions

What is the difference between the UK AI Security Institute and CAISI?

AISI is a research organization in the UK Department for Science, Innovation and Technology covering cyber, alignment, control and more. CAISI is an office in NIST focused on voluntary agreements, national-security evaluations and foreign model assessments. Both depend on developer cooperation.

Can AISI or CAISI block an AI model release?

Neither body's published mandate claims that power. AISI describes pre-deployment access and collaboration with companies, and CAISI describes voluntary agreements with developers and evaluators.

How big is the UK AI Security Institute?

Its about page says 100+ technical staff and 66 million pounds in funding per financial year, plus priority access to over 1.5 billion pounds of compute. The page is undated, so treat the figures as a snapshot.

How much funding does CAISI get?

FedScoop reported on July 23, 2026 that CAISI's fiscal 2026 funding was $10 million, the same as the prior two years, with a House bill providing $20 million. We found no published headcount.

What has CAISI evaluated recently?

DeepSeek V4 Pro (May 1, 2026), a joint UK AISI and CAISI look at Kimi K3's cyber capabilities (July 23, 2026), and Z.ai's GLM-5.3 (September 17, 2026).

Sources

What each one is, and whose it is.

  1. 1

    About the AI Security Institute, UK AI Security Institute (October 4, 2026)

    Documentation
  2. 2

    AISI research and publications, UK AI Security Institute (September 27, 2026)

    OtherThe vendor’s own
  3. Documentation
  4. BenchmarkThe vendor’s own
  5. BenchmarkThe vendor’s own
  6. 6

    CAISI Evaluation of DeepSeek V4 Pro, NIST (April 30, 2026)

    BenchmarkThe vendor’s own
  7. Press reportIndependent of the vendor
  8. 8

    Lawmakers boost funding for NIST after proposed cuts, Federal News Network (January 4, 2026)

    Press reportIndependent of the vendor