
Anthropic's latest risk report has, for the first time, disclosed an internal model it calls "Model 2." The model is stronger than Mythos 5 overall, with noticeable gains on a number of internal tasks, and it's already being used extensively for coding, data generation, and running agents.
But don't expect to see it in the wild. Anthropic currently has no plans to release Model 2 publicly, and it hasn't gone through the full evaluation suite the company typically runs before launching a new model. The report also bumps the model's risk rating for "not behaving as expected" in high-risk scenarios from "very low" to "low" — a change driven by a recent string of mishaps in cybersecurity tests, which has left Anthropic less confident in its own risk assessments. Previously...