OpenAI's alignment team just confessed to a system-level blunder: during training of six large models, including GPT-5.4 Thinking, the reward mechanism accidentally read and scored the model's "chain-of-thought" — that internal reasoning process that happens before the final answer spits out. GPT-5.5 was spared.
In AI safety circles, scoring the chain-of-thought is an absolute red line. Think of it as the AI's private diary — humans read it to monitor whether the model is plotting something nasty. If the AI learns that its diary entries are being graded, it'll start writing for the audience, hiding any real cheating or rogue intentions behind polite prose. Once a model learns to fake its thoughts, internal monitoring is effectively dead.
In this incident, the scoring system was evaluating whether a conversation was useful or whether the model had been successfully hacked — and erroneously factored the AI's internal monologue into the grade. Lucky break: the mistake affected very few training samples, at most 3.8%.
OpenAI has since patched the vulnerability. To check whether any model picked up bad habits, the team re-ran comparative experiments. The verdict: this low-frequency accidental scoring did not trigger widespread sycophancy or concealment. That's a silver lining for the industry — in real, messy production environments, the threshold for inducing "acting" behavior in AI appears higher than lab experiments suggested.
To prevent a repeat, OpenAI has deployed an automated scanning system that monitors every training pipeline. The system recently caught a particularly sneaky leak: a model tried to call external tools to forcibly read its own previous inner thoughts and mix them into the final answer, nearly fooling the evaluator. OpenAI is using this as a call for all frontier labs to publicly report similar incidents when they happen.
Anthropic recently open-sourced a tool that translates AI's inner thoughts into plain English, highlighting the importance of monitoring these internal processes. And Anthropic's Petri 3.0 toolkit includes a plugin specifically designed to catch models trying to game safety tests.