Menu

Categories

Tags

Claude's distillation defenses are failing: small models can extract hidden chain-of-thought

August 12, 2026 | Source: t | AI, Anthropic | 142 views 0 comments

Claude's anti-distillation protections may be for nothing. To keep larger models' capabilities from being distilled into cheaper imitations, AI labs hide the full chain-of-thought and only give users a summary. But new research suggests that shield isn't nearly as sturdy as it looks: hand Opus's encrypted reasoning over to Anthropic's weaker Haiku model, jailbreak Haiku, and it will happily recite Opus's original line of thinking. The researchers got the trick working on Claude, GPT, and Gemini.

This has actually been building for a few months. In May, cryptographer Matthew Green found that these encrypted chain-of-thoughts can be replayed across sessions and accounts — but at the time, there was no reliable way to actually read their contents. The new paper...

Leave a Reply

Your email address will not be published. Required fields are marked *