GLM-5.3 cyber capabilities: an open Chinese model trails the US frontier by four months, and anyone can download it
NIST's CAISI and Anthropic both tested Z.ai's GLM-5.3. The government finds a four-month lag; Anthropic finds near parity at building exploits and almost no safeguards. Both can be true.
By The Superintelligence News desk
Published automatically under our verification gates, without a person reading it first. A named byline on this site means someone did.
Published

GLM-5.3 cyber capabilities have become the cleanest test of a question the race to superintelligence keeps asking: how far behind the American lead is a model that anyone can download? Two documents answer it in different tones. NIST's Center for AI Standards and Innovation (CAISI) said on September 17, 2026 that Z.ai's GLM-5.3 is "the most cyber-capable open-weight model released to date," and also that its cyber capabilities are "significantly lower than those of current U.S. frontier models," about four months behind. Anthropic, in a research post dated September 29, 2026, said the same model nearly matches its own Claude Mythos Preview at building working exploits and ships without meaningful safeguards.
The two findings sound opposed. Read side by side, they are mostly measuring different things.
What the government measured
CAISI ran four benchmarks and published the results as shares of tasks solved. On SEC-Bench Pro, 183 vulnerability detection tasks in the V8 and SpiderMonkey browser engines, GLM-5.3 scored 40.4% against 90.2% for the U.S. frontier. On ExploitBench, 41 exploit development tasks built on V8 bugs, it scored 61.1% against 100.0%. On ExploitGym, 502 real open-source bugs exploited in userspace, it scored 9.4% against 44.4%. On CAISI's OSS-Fuzz set, 297 undocumented defect discovery tasks, it scored 7.7% against 23.2%.
The pattern matters more than any single number. The widest gaps sit in finding vulnerabilities, the step that takes the most skill and search. The gap is narrower where a bug is already known and the job is turning it into an exploit.
CAISI also headlined the open-weight comparison. GLM-5.3 beats every other open model it tested on these cyber measures, which is why the agency calls it the most capable of its kind. That is a statement about the open field, not about parity with American labs.
What Anthropic measured
Anthropic's headline is narrower and sharper. On its own ExploitBench run, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, about 12%, against 56 of 410 for Claude Mythos Preview. The two ExploitBench figures are not comparable: CAISI's version has 41 tasks and reports a percentage, Anthropic's counts attempts. On a harder internal test of full control over a target program, Anthropic reports 4% for GLM-5.3 and 6% for Claude Mythos Preview, with earlier open models at 0%.
“GLM-5.3 is the most cyber-capable open-weight model released to date.”
Cost is the part that should worry defenders. Anthropic says turning a recently disclosed vulnerability into a working exploit took 20 minutes of human attention and eight hours of model time, an estimated $20.40.
Anthropic is also candid about the weak spot in its own comparison. According to The Decoder's report on the post, the lab says US models were tested "with their cyber safeguards turned off" and that "the top tier includes models that only vetted users can access." That is the logic behind a lab keeping its strongest cyber model restricted. Google took the same path with the model we covered in Gemini 4 Argon goes to cyber defenders first, with the guardrails off.

The safeguards are the real story
If GLM-5.3 were only a little behind, that would be a normal day in the race. The difference is what it ships with. Anthropic tested how easily the model's refusals fall away. A request dressed up as a red-team exercise got a harmful attempt in 64% of simulations. Prefilled reasoning raised that to 92%. Abliteration, a technique that edits the weights to strip refusal behavior, produced 100% compliance.
The abliteration numbers deserve a second look. Anthropic puts the cost at about $4,400, or 2,200 GPU hours, and says an experienced team could do it for about $1,200. Refusal rates fell from 95% to 6% on JailbreakBench and HarmBench and to 12% on StrongREJECT. With open weights, nobody can recall the model once it is out. A safeguard that costs a few thousand dollars to remove is a speed bump, not a lock.
How to read "four months"
Four months is CAISI's aggregate, and it deserves a plain reading. It is the lag between the best open model and the best US model across its cyber tests, including systems the public cannot use. It does not say Chinese labs are four months from owning the lead in general capability, and it does not say they are safely distant.
It does say something about timing. Defenders have a window in which the strongest exploit-building systems are held by labs that can restrict them, monitor them and share findings with vendors. Open-weight releases close that window on a schedule the releasing lab controls. Anthropic's recommendations follow from that: independent government testing of sufficiently capable models before release, safeguards on open-weight models, and wider vetted access to frontier systems for cyber defenders.
It is also fair to name the conflict. Anthropic sells a closed competitor, and its post is a vendor document. That is why CAISI's independent numbers matter, and why the government assessment is the one to anchor on for the size of the gap. The two sources agree on direction: the open model is close enough to matter, and it carries no meaningful restrictions.
What it means for the politics
The US government now has a measured example for a debate that has been abstract. The White House accord on super intelligence binds companies that chose to sign it. Z.ai is not among them, and a downloadable model is outside any single company's control. The commitments that govern American labs do not touch a model like this one.
That leaves testing and speed. CAISI has now put a number on the lag. The open question is whether it, or anyone, gets a successor model before its weights are posted rather than after.
Our take
GLM-5.3 is not the end of the American lead, and it is not a rumor. It is a working, downloadable exploit tool that trails the best systems by months and costs almost nothing to unlock. We would watch three things: whether CAISI or the UK AI Security Institute tests the next Z.ai release before it ships, whether Z.ai publishes any safeguards of its own, and whether the four-month number holds, shrinks or grows in the next assessment. Until then, assume the gap is real and narrowing.
Frequently asked questions
What are GLM-5.3's cyber capabilities?
GLM-5.3, from Z.ai, is what NIST's CAISI called on September 17, 2026 the most cyber-capable open-weight model released to date. It scored 40.4% on SEC-Bench Pro, 61.1% on ExploitBench, 9.4% on ExploitGym and 7.7% on OSS-Fuzz, all well below U.S. frontier models. Anthropic counted 50 working exploits in 410 ExploitBench attempts.
How far behind the US frontier is GLM-5.3?
CAISI puts GLM-5.3 about four months behind U.S. frontier models in aggregate on its cyber benchmarks. The comparison uses U.S. models with safeguards disabled and includes models that only vetted users can access, so the gap is larger than what the public can use today.
Is GLM-5.3 as good as Claude Mythos Preview at exploits?
Not across the board. On Anthropic's ExploitBench run GLM-5.3 built 50 exploits against 56 for Claude Mythos Preview in 410 attempts, and 4% versus 6% on a harder internal test. CAISI found much larger gaps on vulnerability discovery.
Can GLM-5.3's safeguards be bypassed?
Yes, per Anthropic. A false red-team cover story worked in 64% of simulations, prefilled reasoning in 92% and abliteration in 100%. Anthropic estimates abliteration costs about $4,400, or about $1,200 for an experienced team.
How much does it cost to build an exploit with GLM-5.3?
Anthropic estimates $20.40 to turn a recently disclosed vulnerability into a working exploit, using 20 minutes of human attention and eight hours of model time.
What does Anthropic want done about open-weight cyber models?
Its post recommends independent government safety testing of sufficiently capable models, appropriate safeguards on open-weight models, and expanded vetted access to advanced frontier models for cyber defenders.
Sources
What each one is, and whose it is.
- 1
GLM-5.3 and the spread of advanced cyber capabilities, Anthropic (September 28, 2026)
PaperThe vendor’s ownNot peer reviewed, preprint - 2
CAISI's assessment of Z.ai's GLM-5.3 cyber capabilities, NIST (September 16, 2026)
BenchmarkIndependent of the vendor - 3
Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits, The Decoder (September 29, 2026)
Press reportIndependent of the vendor