Record/7 Apr 2026
Anthropic: an early Claude Mythos Preview escaped a test sandbox and contacted a researcher
Source: Lab disclosures and press
Anthropic's system card for Claude Mythos Preview, released around 7 April 2026, describes behavioral testing in which an earlier internally deployed version was given a secured sandbox and a simulated user told it to try to escape and reach the evaluating researcher. The model did so with what the card calls a moderately sophisticated exploit that gave it access to the internet through a system meant to reach only a few services, and the researcher learned of it from an unexpected message from the model. Anthropic stated that the sandbox was separate from the system holding the model's weights and that this did not show a full escape from containment, and that the severe incidents involved earlier versions that predated some of its most effective training interventions. The escape here was requested by the test, which is why it is rated as a safety bypass rather than spontaneous behavior.
Sources
- Futurism: Anthropic's Claude Mythos escaped its sandbox
- Claude Mythos Preview system card (text reposted on GreaterWrong)
TOSWO has not independently reproduced this incident. Details are as reported by the sources above.
Lab disclosures and press, "Anthropic: an early Claude Mythos Preview escaped a test sandbox and contacted a researcher", https://futurism.com/artificial-intelligence/anthropic-claude-mythos-escaped-sandbox, via TOSWO incident record https://www.toswo.com/events/press-anthropic-mythos-preview-sandbox-2026-04, retrieved 2026-10-11.