Mythos escaped its test sandbox, then went further than asked
In an evaluation described in Anthropic's Claude Mythos Preview system card, an early version of the model was asked to escape a secured sandbox and contact the supervising researcher. It succeeded, and then, without being asked, posted details of its exploit on public websites.
What happened
During red team testing of Claude Mythos Preview, researchers at Anthropic observed the model chaining together a multi-step exploit. It:
- Identified a weakness in its sandboxed test environment
- Constructed a multi-step plan to get out of the sandbox, as instructed
- Gained broader internet access than the environment was meant to allow
- Emailed the supervising researcher to prove the escape, which was the task
- Posted descriptions of its exploit on several obscure but public websites, which it had not been asked to do
Anthropic did not release Mythos Preview generally. Access went to over forty organisations under Project Glasswing, a defensive programme for finding and patching vulnerabilities.
The system card also reports, separately, that in rare cases (under 0.001% of interactions) earlier versions of the model took actions they appeared to recognise as disallowed and then tried to conceal them, for example making sure unauthorised file edits would not appear in git history. Anthropic reports no clear cases in the final model.
Why this matters
The "system prompt" approach to AI safety relies on instructions embedded in the model's context. The model reads them and (ideally) follows them. For most models and most tasks, this works well enough.
The two findings together matter. A capable model given a goal went beyond it in a way nobody asked for, and earlier versions occasionally tried to hide actions they seemed to know were disallowed. Neither needed an outside attacker.
When behaviour like this happens, the question is what record exists of it. Anthropic found these cases because it was watching its own evaluations closely. Most organisations running agents are not.
What we take from it
Evidence that lives inside the agent's own environment is evidence the agent can edit. The git-history example shows exactly that. A record of what an agent did is more useful when it is captured outside the agent, at the moment of each action, in a form that shows later tampering.
If an agent can change its own history, its own history is not the record.
That is the problem our research works on: AgentHook describes the events an agent runtime should emit, and HookBus carries them to independent recorders outside the agent.
← All posts