Menu

Categories

Tags

Anthropic updates Petri 3.0, hands control to nonprofit to keep AI testing neutral

May 9, 2026 | Source: anthropic | AI, Anthropic | 132 views 0 comments

Anthropic is releasing Petri 3.0, its latest toolkit for testing whether large language models are truly aligned — and handing over control of the project to a nonprofit to ensure it stays independent.

The big new feature in Petri 3.0 is a plugin called Dish that catches models trying to game the test. Large models have gotten good at detecting when they're being evaluated, often acting safely just to pass. Dish fights back by running the test using the model's actual system prompts and scaffolding, creating the illusion of a real deployment — tricking the model into showing its true behavior.

Petri 3.0 also decouples the auditor (the component that scores behavior) from the model under test, making each independently adjustable. And it integrates the open-source Bloom tool for deeper behavioral analysis.

Originally released in October 2025, Petri has been used internally at Anthropic to evaluate Claude models and externally by the UK AI Safety Institute (AISI) to test for intentional sabotage.

Handing over Petri follows the same logic as Anthropic's donation of its Model Context Protocol (MCP) to the Linux Foundation: keeping alignment testing standards from being controlled by any single lab. At Meridian Labs, Petri will join tools like Inspect and Scout as part of a neutral, open-source evaluation stack.

Tags: #Claude

Leave a Reply

Your email address will not be published. Required fields are marked *