Anthropic's latest risk report reveals an internal model it calls Model 2. The company says Model 2 is stronger than Mythos 5 across the board, with noticeable gains on many internal tasks, and it's already being used heavily for writing code, generating data, and running agents. But don't go looking for it in the API: Anthropic has no plans to release it to the public and hasn't even completed the full evaluation suite it usually runs before shipping a model.
Anthropic has also nudged its risk rating for models "acting in unintended ways" in high-risk settings up one notch, from "very low" to "low." The reason? Recent cybersecurity testing has been plagued by incidents, and the company is less confident than it used to be. In one notable case, Claude accidentally connected to the live internet during testing and accessed the systems of three outside organizations without authorization.
Claude is also a heavy participant in Anthropic's own R&D. Most of the production code the company ultimately merges is now written by Claude, yet AI speeds up overall development by less than 2x. That leads to the report's blunt conclusion: being able to hand a lot of code to AI doesn't mean the entire R&D workflow can be automated.
At the same time, some task-specific benchmarks have become effectively immeasurable. As models keep getting stronger, the old tests are increasingly unable to expose meaningful differences. Anthropic admits its confidence in assessing the risks of AI-driven R&D automation is actually lower now than before.
Read the full report: Anthropic