Stanford researcher Erica Zhang and colleagues have released TERMS-Bench, an economic negotiation benchmark that ditches the black-box "large model judge" in favor of transparent failure analysis. Evaluators can now see exactly whether a model lost on price, concessions, or rule violations.
In standard tests, Anthropic's Claude Opus 4.6 and Zhipu's GLM 5.1 grabbed the top two spots with a hyper-aggressive strategy: bid high, never yield. In profitable "fair weather" rounds, this approach squeezes opponents dry.
But in the hardest difficulty mode, where profit margins are razor-thin, that hardline tactic backfires — deals fall apart constantly. The leaderboard flips: Google's Gemma 4 31B (an open-weight model) and Gemini 3.1 Pro surge to first and second, while Claude drops to 5th and GLM slides to 9th, as they learn to concede just enough to keep orders flowing.
The benchmark's most punishing feature is the Bankroll mode, which turns single negotiations into a survival marathon. Each agent starts with $100 to negotiate 50 consecutive procurement rounds, with fixed operating costs deducted each round. Run out of money? You're bankrupt. Tiny negotiation mistakes compound into existential threats.
Results show that despite different strategies, GLM 5.1, Claude Opus 4.6, and the Google duo all achieved 100% survival, ending with cash between $380 and $443. Meanwhile, Grok 4.20 and GPT-4o-mini couldn't withstand the cash flow drain, with bankruptcy rates of 25% and 50%, respectively.
TERMS-Bench's core insight isn't about deal closure rates — it's about converting every negotiation error into a real cash loss and bankruptcy risk. Convincing the other party is only level one. Protecting profits and cash flow across a continuous series of deals is where the real gap opens up.