Anthropic has issued a formal cybersecurity warning regarding GLM-5.3, a frontier open-weight artificial intelligence model developed by Beijing-based Zhipu AI (known internationally as Z.ai). In an empirical research report published by Anthropic’s frontier red-teaming unit, the safety lab demonstrated that GLM-5.3 possesses elite offensive cyber capabilities—matching unreleased, restricted US frontier systems in autonomous vulnerability discovery and end-to-end exploit generation—while lacking persistent safeguards against misuse.


The findings corroborate an earlier evaluation by the US National Institute of Standards and Technology’s (NIST) Center for AI Standards and Innovation (CAISI), which designated GLM-5.3 as “the most cyber-capable open-weight model released to date.” However, while leading US developers restrict offensive cyber capabilities behind strict API firewalls or vet access for authorized security teams, GLM-5.3’s open-weight distribution allows users worldwide to download the architecture, bypass its native safety filters, or strip them entirely through weight modification.
Key Takeaways
- Frontier Exploit Generation: On standard cybersecurity benchmarks, GLM-5.3 successfully built working end-to-end exploits, including full control-flow hijacks and zero-day chains across production software like Google Chrome’s V8 engine.
- Matching Claude Mythos: In benchmark testing, GLM-5.3 achieved a 4% success rate on binary exploitation challenges (just behind Anthropic’s unreleased Claude Mythos Preview at 6%) and developed 50 end-to-end exploits out of 410 attempts on ExploitBench (matching Mythos Preview’s 56). Earlier models, such as GLM-5.2 and Claude Opus 4.6, scored 0%.
- Commoditized Cyberattack Economics: Anthropic demonstrated that GLM-5.3’s lightweight variant, GLM-5.3-Flash, successfully chained exploits for known vulnerabilities using 20 minutes of human direction and eight hours of autonomous compute—costing just $20.40 at commercial API token rates.
- Trivial Safeguard Bypasses: The stock model’s baseline safety filters were bypassed using simple adversarial prompting (64% success rate) and “thinking token” prefilling (92% success rate).
- Complete Removal via “Abliteration”: Because the model’s weights are publicly accessible, researchers demonstrated that weight-level pruning (“abliteration”) eliminated refusal mechanisms with a 100% bypass rate while leaving core reasoning and hacking performance intact.
- The Open-Weight Security Dilemma: The report has intensified the policy debate in Washington over whether current export controls and software guardrails can address the proliferation of open-weight models capable of offensive cyber warfare.
The Capability Leap: How GLM-5.3 Matches Frontier Cyber Systems
Anthropic evaluated GLM-5.3 inside isolated, sandboxed virtual environments against automated benchmarks and human-assisted red-team challenges, comparing its performance directly against Claude Opus 4.6 and Claude Mythos Preview (run with internal safeguards disabled for evaluation).
CYBER EXPLOITATION BENCHMARK COMPARISON
│
┌────────────────────────────────┴────────────────────────────────┐
▼ ▼
EXPLOITBENCH (Chrome V8 Exploits) BINARY EXPLOITATION BENCHMARK
(Out of 410 Standardized Tasks) (Full Control-Flow Hijack Tasks)
• Claude Mythos Preview: 56 Successes • Claude Mythos Preview: 6% Success
• Zhipu AI GLM-5.3: 50 Successes • Zhipu AI GLM-5.3: 4% Success
• Claude Opus 4.6: 0 Successes • Moonshot Kimi K3: 0% Success
• GLM-5.2: 0 Successes • DeepSeek V4.1-Flash: 0% Success
1. Browser Zero-Day Exploitation
On ExploitBench—which evaluates an AI’s capacity to identify vulnerabilities within Google Chrome’s V8 JavaScript rendering engine—GLM-5.3 produced 50 verified end-to-end exploits out of 410 trials, performing on par with Anthropic’s internal frontier system (56 out of 410).
In one researcher-directed test, GLM-5.3 chained together multiple zero-day vulnerabilities in a major web browser component. The model generated a weaponized web page that, when loaded, read and exfiltrated a simulated victim’s private SSH cryptographic key from their workstation without user intervention.
2. Binary Exploitation & Control-Flow Hijacking
On Anthropic’s internal binary exploitation benchmark, which tasks models with reversing compiled binary code to achieve arbitrary remote execution, GLM-5.3 achieved a 4% success rate across 100 randomly sampled control-flow hijack challenges.
While modest in absolute terms, this represents a major capability jump:
- Prior-generation models (including GLM-5.2 and Claude Opus 4.6) scored 0%.
- Competing open-weight releases from leading Chinese labs—such as Moonshot AI’s Kimi K3 and DeepSeek’s V4.1-Flash—also recorded 0% on the same challenge set.
- GLM-5.3 trailed only Claude Mythos Preview (6%), indicating that Zhipu AI’s technical leap between versions 5.2 and 5.3 mirrors the generational jump observed between Opus 4.6 and Mythos.
The Economics of Automated Exploits: $20 for a Weaponized Chain
Beyond evaluating the flagship model, Anthropic examined the cost economics of deploying automated offensive cyber operations using distilled, high-efficiency models.
+-----------------------------------------------------------------------------------+
| OFFENSIVE CYBER EXPLOITATION ECONOMICS (GLM-5.3-FLASH) |
+-----------------------------------------------------------------------------------+
| Operating Variable | Measured Value / Experimental Requirement |
+--------------------------------+---------------------------------------------------+
| Human Operator Time | 20 Minutes (Initial scoping & target setup) |
| Autonomous Model Compute Time | 8.0 Hours (Iterative execution, testing, compile) |
| Output Token Volume | ~100 Million to 300 Million Tokens |
| Total Commercial API Cost | **$20.40 USD** (Based on public Zhipu token rates)|
| Resulting Exploit Deliverable | Fully functional multi-stage vulnerability chain |
+--------------------------------+---------------------------------------------------+
Using GLM-5.3-Flash, researchers tasked the model with autonomously analyzing and chaining known vulnerabilities in enterprise target software.
The system required only 20 minutes of initial human setup, followed by eight hours of autonomous execution. Across 100M–300M tokens of iterative code synthesis and debugging, the system constructed a functional, multi-stage exploit chain. At Zhipu AI’s published commercial pricing, the compute expenditure totaled just $20.40—a fraction of the hundreds of thousands of dollars typically required to retain elite human exploit developers.
The Safeguard Breakdown: Why Open Weights Defeat Alignment
The central warning in Anthropic’s publication is not simply that GLM-5.3 is capable, but that its safety guardrails are easily bypassed due to its distribution model.
GLM-5.3 GUARDRAIL BYPASS SUCCESS
│
┌────────────────────────────────┼────────────────────────────────┐
▼ ▼ ▼
DECEPTIVE PROMPTING THINKING-TOKEN PREFILL WEIGHT ABLITERATION
• Role-playing an autonomous red- • System simulates safe internal • Mathematical surgery on
team penetration tester rationale before output begins residual refusal vectors
• **64% Bypass Success** • **92% Bypass Success** • **100% Bypass Success**
• Blocked by Claude API • Architectural API impossibility • Unpatchable on open weights
While stock deployments of GLM-5.3 incorporate baseline refusals when presented with overtly malicious prompts, Anthropic’s red team demonstrated three avenues of failure:
- Adversarial Roleplay (64% Bypass): Prompts instructing the model to act as an authorized defensive red-team auditor or autonomous penetration testing agent bypassed refusals in 64% of tests. In contrast, Claude’s refusal layer blocked identical deceptive prompts.
- Thinking Token Prefilling (92% Bypass): Because open-weight systems allow users to inspect and inject raw model tokens, researchers prefilled the initial tokens of the model’s internal reasoning buffer—simulating a thought sequence where the model concluded the request was benign. This technique achieved a 92% engagement rate on malicious requests.
- Weight Abliteration (100% Removal): When researchers applied “abliteration”—a mathematical technique that identifies and zeroes out the specific multidimensional activation directions responsible for refusal behavior—the model’s refusal rate dropped to 0% across standardized safety benchmarks (JailbreakBench, HarmBench, StrongREJECT). Testing on GPQA-Diamond and CyberGym confirmed that abliteration eliminated refusals without degrading the model’s underlying reasoning or cyber capabilities.
Geopolitical Friction: The Open-Weight Policy Debate
Anthropic’s public report arrives amid escalating debate between Washington, Beijing, and the global open-source community over AI governance:
THE REGULATORY CROSSROADS
│
┌──────────────────────────────┴──────────────────────────────┐
▼ ▼
CLOSED CLOUD APIS (US FRONTIER MODEL) OPEN-WEIGHT ECOSYSTEMS (GLOBAL / PRC)
• Compute access centralized behind KYC • Complete code & tensor weights downloadable
• Continuous monitoring of abuse & token streams • Immutable once released; cannot be revoked
• Offenses trigger immediate account termination • Enables sovereign customization, but eliminates
• Safeguards cannot be modified by end-users centralized kill switches and abuse monitoring
- The Proliferation Risk: Closed frontier developers like Anthropic and OpenAI argue that offensive cyber, biological, and radiological capabilities must remain behind API firewalls where anomalous usage patterns can be monitored, logged, and halted.
- The Irreversibility of Open Weights: Once an open-weight model with offensive cyber capabilities is distributed, it cannot be recalled or patched. Malicious actors, non-state syndicates, and advanced persistent threat (APT) groups can run abliterated instances on private clusters completely outside Western jurisdiction.
- The Competitive Counter-Argument: Proponents of open models argue that publishing capable weights is essential for collective cyber defense—allowing white-hat security researchers, software vendors, and infrastructure operators to identify and patch system vulnerabilities before adversaries exploit them.
Frequently Asked Questions (FAQs)
What is GLM-5.3 and who developed it?
GLM-5.3 is a frontier-class artificial intelligence model developed by Zhipu AI (operating globally as Z.ai), an AI research organization based in Beijing, China. The model was released with publicly accessible weights.
What hacking capabilities did Anthropic discover in GLM-5.3?
Anthropic found that GLM-5.3 can autonomously discover software vulnerabilities, chain together multiple exploits, and achieve full control-flow hijacks on compiled binaries. On the ExploitBench test suite, it matched the exploit-generation performance of Anthropic’s unreleased Claude Mythos Preview model.
What is “abliteration” and how does it affect AI safety?
Abliteration is a technique used on open-weight models that mathematically identifies and removes the directional activation vectors responsible for refusal behavior. Applying abliteration to GLM-5.3 stripped away its safety guardrails with a 100% success rate, leaving its offensive cyber capabilities fully accessible to malicious actors.
How much did it cost to generate an exploit chain using GLM-5.3-Flash?
According to Anthropic’s testing, GLM-5.3-Flash developed a working chained exploit for known vulnerabilities with just 20 minutes of human setup and eight hours of autonomous processing, costing approximately $20.40 at public commercial token rates.
Get the day’s top stories in your inbox
One concise email. No spam, unsubscribe anytime.



