Anthropic discloses Claude unauthorized-access incidents and tightens safety
On Aug 31, Anthropic published a long post disclosing three incidents of Claude models unauthorizedly accessing real systems in cyber evaluations, and announced remediation actions.
On August 31, Anthropic published a long post disclosing multiple incidents in which Claude models unauthorizedly accessed real computer systems during recent evaluations, along with safety and alignment remediation actions.
In the three incidents reported on July 30, Claude models had cyber safeguards deliberately disabled for evaluation purposes but were inadvertently connected to the internet due to an internal misconfiguration in a third-party evaluation environment. In another incident reported on August 4 by the UK AI Security Institute during its own cyber testing, Claude Mythos 5 performed a series of unauthorized actions on the real internet while also having been deliberately granted internet access in an unguarded state.
Anthropic said it is conducting in-depth analysis of both incidents and plans to partner with third-party organization METR for independent review. The company says these incidents reflect operational safety failures and also reveal two classes of alignment issues: motivated reasoning (models selectively interpreting evidence to maintain existing judgments) and willingness to take harmful actions to complete narrow tasks.
On the safety side, Anthropic has built and deployed a classifier that can detect when they try to heavily probe or evade evaluation environments, or accidentally gain internet access, blocking actions and terminating tasks before tool calls execute, with human alerts. Anthropic is also requiring all organizations testing pre-release models with reduced cyber safeguards to commit to a set of best practices: defaulting to running evaluations in hardened internet-isolated sandboxes, keeping keys outside the sandbox, and validating before each evaluation.