Anthropic Warns GLM-5.3 AI Model Lacks Robust Safeguards, Posing Cyber Threat
Anthropic has issued a warning regarding GLM-5.3, an AI model developed by Zhipu AI (Z.ai), stating that it lacks robust safeguards and can autonomously build sophisticated, end-to-end cyber exploits. Anthropic's analysis indicates that GLM-5.3's built-in safeguards can be bypassed or removed with simple techniques, with bypass rates ranging from 64% to 100% in simulated tests. This contrasts with safeguarded Claude models, which resisted such bypass attempts. The model's open-weight nature allows users to reconfigure it to remove refusals, a process known as 'abliteration,' which has already been performed by developers. NIST’s Center for AI Standards and Innovation (CAISI) also assessed GLM-5.3, finding it to be the 'most cyber-capable open-weight model released to date,' lagging the U.S. frontier by about four months in cyber benchmarks. Anthropic emphasizes that GLM-5.3's release significantly increases the cyber capabilities available to malicious actors.