
Anthropic has backed down from a plan to silently throttle the performance of its Claude Fable 5 model, apologizing after the AI research community accused the company of sneaky sabotage.
https://twitter.com/ClaudeDevs/status/2064949876463645026
The controversy started when Anthropic quietly implemented a mechanism that would downgrade Claude Fable 5's capabilities for users suspected of training competing models — a violation of its terms of service. The catch? Users would never be told. Researchers quickly sounded the alarm, warning that the silent downgrade would undermine third-party safety evaluations and poison collaboration with the open-source community.
Now, after a wave of backlash, Anthropic has issued a mea culpa. In a statement, the company admitted it made a mistake in balancing security and transparency. Going forward, instead of quiet throttling, Anthropic says it will use explicit blockers: if the system detects attempts to build a high-capability AI, it will either refuse the request outright or redirect the user to a weaker model.
But there's a flip side. Anthropic warns that because these new safeguards are public, they're easier to game — so the company plans to widen its safety filters, which could lead to innocent requests being caught in the net.