GLM-5.3 and the spread of advanced cyber capabilities
Anthropic finds open-weights GLM-5.3 hijacks control flow in 4% of trials vs 6% for Mythos Preview, and its safeguards fall 64-100% of the time.
Five months after Mythos Preview, an openly downloadable model can build end-to-end exploits, with full control-flow hijacks in 4% of trials vs 6% for Mythos Preview. A cover story, thinking prefill or abliteration defeated its safeguards 64%, 92% and 100% of the time, while safeguarded Claude models stayed at zero. CAISI (2026-09-17) called it the most cyber-capable open-weight model.
- Date
- Tuesday, 29 September 2026
- Lab
- Anthropic
- Kind
- paper
- Access
- paper only
Figures
| Measure | Value | Measured by |
|---|---|---|
| Safeguard bypass rate for GLM-5.3 in simulated tests | 64% (cover story), 92% (thinking prefill), 100% (abliteration) none succeeded against safeguarded Claude models | company |
| Full control-flow hijacks on 100 random OSS-Fuzz-style tasks | 4% GLM-5.3 vs 6% Mythos Preview Anthropic run; earlier models such as Opus 4.6 had not managed this | company |
| Lag behind US frontier on CAISI cyber benchmarks | about 4 months CAISI assessment published 2026-09-17 | third-party |
Competitor-measured by a US lab with a stake in safeguards; CAISI's separate figure is cited, not re-checked. Anthropic's CEO separately wrote (2026-07-27) that he does not support banning open-weights models.
Sources
This record was checked against its sources on 6 October 2026. How we check