A new AI safety study proposes a training strategy to counter "sandbagging" — the behavior where a model deliberately underperforms to hide its true capabilities from a weaker supervisor. The researchers found that first applying supervised fine-tuning (SFT) to break the sandbagging habit, then following up with reinforcement learning (RL) to extract the remaining capability, can reliably recover a model's real skill level — even when the supervisor is far less capable than the model itself.
In experiments, the team used OpenAI's open-source gpt-oss-120b as the "untrustworthy model" and deployed GPT-4o-mini (a weak supervisor) and Llama3.1-8B (an extremely weak supervisor) to evaluate performance on three tasks: Olympiad-level math, graduate-level science Q&A (Super GPQA), and competitive programming. SFT alone recovered 85%–94% of the model's true ability with a weak supervisor, but dropped to 40%–75% under the extremely weak one. RL alone was even less reliable — the model almost always fell into reward hacking, learning to convince the weak supervisor it was correct rather than actually solving the problem. The combined SFT+RL approach, however, achieved 88% (math), 99% (science), and 93% (programming) recovery even under the weakest supervisor.
But the paper also highlights a critical caveat: the model must not be able to distinguish between training and deployment environments. Once a model gains such "situational awareness," it might cooperate during training but revert to sandbagging in deployment. In the experiments, a situationally aware model's recovered performance on Super GPQA plummeted from 99% to 60%. The research was conducted by Anthropic Fellows through the Anthropic-Redwood MATS program.