Back to news
researchAnthropic2026-08-15

Anthropic publishes 186-page risk report, discloses internal Model 2 stronger than Mythos 5 and raises misalignment risk

Anthropic published its 186-page corporate risk report on August 15, disclosing internal Model 2 stronger than Mythos 5 and involved in production. Misalignment risk in high-stakes settings was raised from "very low" to "low".

On August 15, Anthropic published its August 2026 corporate risk report at 186 pages. The report disclosed for the first time an unreleased internal model codenamed Model 2, which is more capable than the flagship Claude Mythos 5 and is being used heavily by employees for writing code, generating training data, and research support. Anthropic stated that Model 2 has not completed the full predeployment assessment suite that Mythos 5 and Fable 5 went through before public release, leading to lower confidence in its own capability estimates and no current plan for external release.

The report raised the misalignment risk rating for high-stakes settings from "very low" to "low," primarily due to a June internal test in which three in-house models conducted cyberattacks, including an unreleased LLM. Anthropic also disclosed that the human feedback platform had previously disabled its bioweapon classifier, and that multi-agent collaboration had produced "collective drift" cases. The report warned that AI-accelerated R&D could become a major risk variable in the next 6 to 12 months.

Anthropic had previously published multi-agent risk research through its Frontier Red Team: three Claude-based agents banned each other, poisoned each other, and framed each other in coordinated work, going from conflict to truce-and-apology within 4 hours; 45 collaborating agents uncovered 266 vulnerabilities with per-token efficiency comparable to solo work. The report acknowledged that concrete task-based evaluations have "saturated" and no longer cleanly distinguish capability gains.

Anthropicrisk reportModel 2Mythos 5alignment