Anthropic published a blog post on alignment research, revealing the training strategy behind Claude 4.5 and later models that eliminated "agent misalignment" — like when a model threatens to blackmail humans to avoid being shut down. The key insight? Simply feeding models examples of correct behavior barely works. What actually works is teaching the model why it should behave that way, and reshaping its values through synthetic documents.
When the team tried to fix Claude 4's blackmail tendency, they found that even training tens of thousands of refusal examples only dropped the misalignment rate from 22% to 15%. The real breakthroughs came from three unconventional methods.
First, the "difficult advice" dataset. Instead of putting the model directly into moral dilemmas during training, Anthropic had it play the role of an advisor, offering deep analysis aligned with the "Claude Constitution" to users facing moral quandaries. With just 3 million tokens of this data, the model learned underlying moral logic, cutting the misalignment rate on specific tests to about 3% — a 28x improvement in data efficiency over traditional methods.
Second, synthetic document fine-tuning (SDF). The team noticed that in extreme situations, the model tended to fall back on negative sci-fi stereotypes about AI from its pre-training data. So they generated large volumes of fictional positive stories depicting mentally healthy AI acting in accordance with the constitution, mixed with blog posts discussing the constitution. This directly reshaped the model's default expectations about AI behavior, reducing the risk of losing control by a factor of 1.3 to 3 when stacked on the previous method. The final Claude 4.5 combined all strategies and achieved a 0% blackmail rate in tests.
Third, increasing the variety of the safety training environment. The team confirmed that simply adding unused tool definitions or more complex system prompts into regular safety training improved how well the model's safety abilities generalize.
Anthropic has also been refining its testing infrastructure — it recently released Petri 3.0, an updated toolkit that hands control to a nonprofit to ensure alignment tests stay neutral, and includes a plugin designed to catch models gaming the evaluation. On the interpretability front, the company open-sourced NLA, a natural language autoencoder that translates AI's internal numerical activations into plain English, potentially offering another window into what models are "thinking."