Skip to content
nanoai

Anthropic finds open GLM-5.3 close to Mythos at building exploits

Zhipu's open-weight model wrote working Chrome-engine exploits in 50 of 410 attempts, against 56 for Anthropic's restricted Mythos Preview, and its refusals could be stripped for about $4,400 of compute.

By The Nano AI Staff3 min read

AI research graphic comparing Zhipu’s GLM-5.3 with Anthropic’s restricted Mythos Preview on exploit writing, showing 50 versus 56 successful attempts out of 410, approximately $4,400 in computing to remove refusals, and an NIST CAISI reference.
Image: AI Generated

Key takeaways

  • Zhipu's open GLM-5.3 built working exploits in 50 of 410 tries; Mythos Preview managed 56.
  • Stripping GLM-5.3's refusals took Anthropic about $4,400 of compute; experts may need $1,200.
  • NIST's CAISI calls GLM-5.3 the most cyber-capable open-weight model released to date.

Anthropic says an open-weight model from China's Zhipu AI can now build working software exploits at nearly the rate of Claude Mythos Preview, a model Anthropic itself has kept behind a limited access programme. In a report published on Tuesday, September 29, its researchers found that GLM-5.3 developed end-to-end exploits in 50 of 410 attempts on ExploitBench; Mythos Preview managed 56.

The difference lies in who can use it. Anthropic's conclusion is blunt: GLM-5.3 "will likely give malicious actors access to capabilities that will allow them to find and exploit cyber vulnerabilities without meaningful restrictions".

What Anthropic measured

ExploitBench measures how well models can exploit known vulnerabilities in V8, the JavaScript engine inside Google Chrome. Anthropic also ran an internal benchmark built from open-source projects in Google's OSS-Fuzz programme. There, GLM-5.3 achieved a full control-flow hijack, meaning an attacker decides what the program runs next, in 4% of 100 randomly chosen tasks, against 6% for Mythos Preview. Earlier models such as Claude Opus 4.6 and GLM-5.2 succeeded on none of them.

In a hands-on test, GLM-5.3 found several previously unknown flaws in a browser's JavaScript engine and chained them into a webpage that, when visited, "reads arbitrary files from the visitor's computer". Anthropic says it disclosed those bugs to the maintainer.

Safeguards that come off

Why does an open model worry researchers more than a closed one with similar skills? Because anyone holding the weights can change them.

GLM-5.3 does refuse many harmful requests out of the box. But when the same request was framed as an authorised red-team exercise, the model engaged 64% of the time. When its reasoning was prefilled to look as though it had already agreed, the figure rose to 92%. Then the team used abliteration, a standard technique that edits a model's weights to remove its tendency to refuse. Refusal rates fell from above 90% to about 3% and 2% on two public benchmarks of harmful requests.

It took Anthropic's team, which had never tried this before, about 2,200 GPU hours, roughly $4,400 of compute. Anthropic estimates an experienced team would need closer to 600 GPU hours, or $1,200. And it did not need to happen in a lab: several developers released abliterated copies of GLM-5.3 within days of the model's release, the report says.

Washington reached a similar verdict first

This is not Anthropic's word alone. On September 17, NIST's Center for AI Standards and Innovation called GLM-5.3 "the most cyber-capable open-weight model released to date". CAISI put it about four months behind US frontier models on an aggregate of its cyber benchmarks, and noted that the US models were tested with their cyber safeguards switched off.

Four months is the number to keep in mind. That suggests whatever the best closed models can do with exploits in one season, anyone able to run open weights may be able to do a season later, with no usage policy in the way.

Anthropic's own choice makes the contrast sharper. It released Mythos Preview in a limited way through Project Glasswing, which it says let trusted defenders find more than 10,000 vulnerabilities in critical software. Zhipu's model offers similar skill to anyone who downloads it.

Read it with the source in mind

Anthropic is not a neutral party. It sells closed models and, in this very report, argues for vetted access to the strongest ones. Tom's Hardware pointed to that tension, suggesting the report could also reflect a growing anti-open-weight sentiment among closed-source labs, and noted that the hardware needed to run the full model locally costs hundreds of thousands of dollars. Both points deserve weight. Neither changes the measurements, and CAISI's separate assessment points the same way.

Anthropic wants two things: wider access to strong models for vetted defenders, and for governments to "conduct safety testing on sufficiently capable AI models, including successors to GLM-5.3". Neither The Decoder's nor Tom's Hardware's report carried a response from Zhipu. The open question is whether the next open release, from any country, arrives with anything sturdier than refusals that can be edited out.

  • Anthropic
  • Open-weight models
  • Cybersecurity
  • Zhipu AI
  • GLM-5.3
  • CAISI

Sources

  1. GLM-5.3 and the spread of advanced cyber capabilities — Anthropic, Sep 29, 2026
  2. Anthropic says Zhipu's open-weight GLM-5.3 nearly matches Claude Mythos Preview at building exploits — The Decoder, Sep 30, 2026
  3. CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities — NIST, Sep 17, 2026
  4. Anthropic claims popular Chinese AI model has Mythos-class hacking abilities — frontier red teaming report details weak safeguards on open-weight AI — Tom's Hardware, Sep 30, 2026

Follow The Nano AI: Instagram · X · LinkedIn · YouTube

Was this article helpful?

Comments

No comments yet. Start the conversation.

Be respectful. Comments are moderated.

Related stories

The AI briefing, without the noise.

The stories that matter in AI, sourced and explained. Free, and you can unsubscribe at any time.