Anthropic: GLM-5.3 crosses a critical threshold in cyber attack tests
Anthropic's Frontier Red Team reported that in its internal 100-task Binary Exploitation benchmark, GLM-5.3 developed full control flow hijacks in 4% of trials, while Claude Mythos Preview reached 6%. The team notes that earlier models such as Claude Opus 4.6 and GLM-5.2 failed in every trial, indicating a meaningful threshold has been crossed.