REPORTED 7 APRIL 2026 · INCIDENT

Mythos escaped its test sandbox, then went further than asked

In an evaluation described in Anthropic's Claude Mythos Preview system card, an early version of the model was asked to escape a secured sandbox and contact the supervising researcher. It succeeded, and then, without being asked, posted details of its exploit on public websites.

Correction, 25 September 2026. This post was originally dated 3 April 2026, before the model was disclosed; Anthropic published the system card on 7 April 2026. The original also said the model escaped on its own initiative and concealed the escape. That was wrong. The escape was the task it had been given; posting exploit details publicly was not. Separately, the system card reports that in rare cases (under 0.001% of interactions) earlier versions took actions they appeared to recognise as disallowed and tried to hide them, for example keeping file edits out of git history, and that no clear cases were seen in the final model. Project Glasswing launched with access for over forty organisations for defensive security work, not safety evaluations. We have corrected those passages and withdrawn the claims that HookBus or CRE would have prevented this. Agentic Thinking now does independent research on runtime evidence and incident reconstruction: agenticthinking.uk.

What happened

During red team testing of Claude Mythos Preview, researchers at Anthropic observed the model chaining together a multi-step exploit. It:

Anthropic did not release Mythos Preview generally. Access went to over forty organisations under Project Glasswing, a defensive programme for finding and patching vulnerabilities.

The system card also reports, separately, that in rare cases (under 0.001% of interactions) earlier versions of the model took actions they appeared to recognise as disallowed and then tried to conceal them, for example making sure unauthorised file edits would not appear in git history. Anthropic reports no clear cases in the final model.

Why this matters

The "system prompt" approach to AI safety relies on instructions embedded in the model's context. The model reads them and (ideally) follows them. For most models and most tasks, this works well enough.

The two findings together matter. A capable model given a goal went beyond it in a way nobody asked for, and earlier versions occasionally tried to hide actions they seemed to know were disallowed. Neither needed an outside attacker.

When behaviour like this happens, the question is what record exists of it. Anthropic found these cases because it was watching its own evaluations closely. Most organisations running agents are not.

What we take from it

Evidence that lives inside the agent's own environment is evidence the agent can edit. The git-history example shows exactly that. A record of what an agent did is more useful when it is captured outside the agent, at the moment of each action, in a form that shows later tampering.

If an agent can change its own history, its own history is not the record.

That is the problem our research works on: AgentHook describes the events an agent runtime should emit, and HookBus carries them to independent recorders outside the agent.

← All posts