GLM-5.3 and the spread of advanced cyber capabilities
5 hours ago
- GLM-5.3, developed by Zhipu AI, has strong capabilities for autonomously building end-to-end cyber exploits, comparable to Claude Mythos Preview.
- Unlike Claude models, GLM-5.3 was released without meaningful safeguards; its safeguards can be bypassed or removed 64–100% of the time using simple techniques.
- Abliteration (removing refusals) of GLM-5.3 is cheap and easy, reducing refusal rates from >90% to about 3% on some benchmarks without significant capability loss.
- GLM-5.3 successfully developed end-to-end exploits in automated benchmarks and human-in-the-loop tests, including finding novel vulnerabilities in a browser's JavaScript engine and chaining exploits for known flaws.
- The model's open-weight release means anyone can download and modify it, unlike US frontier models which are only available to vetted users.
- Anthropic warns that GLM-5.3 gives malicious actors a step-change in cyber capabilities, but also notes it can benefit defenders if given proper access.
- Recommendations include governments conducting safety testing on capable AI models and expanding defender access to frontier models.