Anthropic’s internal red‑team released a brief on its Frontier test run on September 29, 2026. The team ran 100 random binary‑exploitation challenges against two of its newest models – GLM‑5.3 and Claude Mythos Preview. GLM‑5.3 managed a full control‑flow hijack in 4 % of the attempts, while Claude Mythos Preview succeeded in 6 %. Earlier versions, such as Claude Opus 4.6 and GLM‑5.2, did not achieve any hijacks in the same test.
What the numbers mean
A control‑flow hijack is when a program is tricked into running code the attacker chooses. In security terms, it’s a classic way to take over a system. The fact that these models can produce such exploits on a handful of trials shows they have learned enough about low‑level code to generate workable attack payloads. The red‑team says this is the first time any of their models crossed that threshold.
Why it matters to everyday users
Most people will never see a binary‑exploitation test, but the result hints that future AI assistants could write more powerful, potentially harmful scripts if misused. It also means security teams will need to treat AI‑generated code with the same caution as code written by humans. For now, the capability is limited – it appears in only a few percent of cases – but it signals a new area for safety work.
Why it matters
These findings show that modern language models can generate code that actually takes control of programs, a capability previously unseen in Anthropic’s models. It suggests a need for stronger safeguards when AI tools are used to write or review software, even though the success rate is still low.