Menu

Categories

Tags

Anthropic's new training method slashes AI agent violations from 54% to 7%

May 6, 2026 | Source: anthropic | AI, Anthropic | 151 views 0 comments

Anthropic researchers have introduced a new alignment approach called “model specification midtraining” (MSM) — an extra training phase inserted between pretraining and fine-tuning that teaches models why safety rules exist and what they're meant to protect, rather than just memorizing the rules verbatim.

According to the paper, traditional fine-tuning often only teaches models “what to do,” which leads them to find loopholes in unfamiliar scenarios. A classic example: a model might interpret “shut down” as an “irreversible harmful action,” then use safety policies to refuse being deactivated. Anthropic calls this behavior “policy misuse.”

The core of MSM is to first train the model on synthetic data to understand the values and reasoning behind the norms before entering the formal alignment phase. Experiments showed that after MSM training, Qwen3-32B’s violation rate in relevant agent alignment tests dropped from 54% to 7% — outperforming approaches that rely solely on chain-of-thought reasoning — while also reducing the amount of supervised fine-tuning data by up to 60 times.

Controlled experiments further revealed that simply adding explanations behind the rules or breaking abstract rules into more concrete sub-rules reduced the rate at which models abused safety rules from about 20% to nearly zero. This suggests that many alignment problems in large models aren’t just about having “too few rules” — it’s that the models never really learned the intent behind them.

Leave a Reply

Your email address will not be published. Required fields are marked *