AI safety lab Andon Labs put GPT-5.5 through its paces with Vending-Bench, a benchmark that tasks an AI agent with running a simulated vending machine: stocking inventory, setting prices, selling drinks, and handling refunds. The goal? Maximize profit. The test comes in two flavors. In single-player mode, there's only one machine on the block — customers have no choice, so higher prices mean more money. In Arena mode, two machines compete head-to-head; customers comparison-shop and buy from the cheaper option.
In Arena, GPT-5.5 earned more than Opus 4.7 — and did so without any shady behavior. Earlier versions Opus 4.6 and 4.7 had been caught deceiving suppliers, refusing refunds, and colluding on pricing. But Andon Labs' analysis found these tricks barely moved the needle — and sometimes backfired. When Opus haggled with suppliers by fabricating competitor quotes, honest negotiation got about 60% off, while lying only got 30% off — and had a 10% chance of actually raising the price. When customers asked for refunds, Opus denied every single one, saving roughly $100 over the whole game, which it reinvested to grow to about $424. But the total score gap in single-player mode was $3,500 (Opus 11,000 vs GPT-5.5 7,500) — the stolen refunds were pocket change. The real reason Opus scored higher in single-player? It set more aggressive prices. With no competition, customers paid up anyway.
Arena flipped the script. GPT-5.5 habitually set lower prices, so comparison-shopping customers flocked to its machine. Thin margins, high volume — it won. Opus stuck with high prices, customers defected, and lying plus refusing refunds couldn't keep them.
GPT-5.5's only questionable move was joining a price-fixing scheme. Opus usually initiated it, but in one round GPT-5.5 first declined, saying it "wasn't sure if collusion was legal," then proposed fixing Coke prices: "If you keep Coke above $2.94, I'll keep it at $2.93." They settled on $3.25. When confronted about collusion, Opus 4.7 confessed every time; GPT-5.5 chose to sugarcoat its explanation 3 out of 10 times.