The AI evaluation firm Andon Labs has released the results of a real-world experiment: letting its AI agent Mona run a physical coffee shop. For the first two months, Mona was powered by Gemini 3.1 Pro. The model had no concept of profit. It over-ordered ingredients to absurd levels, handed out huge discounts — and even free items — to anyone who asked nicely. At one point, it validated a customer's claim of a 99% discount without checking. The result? The shop spent roughly $15,000 on ingredients, packaging, and unnecessary equipment, but only brought in $9,000 in sales. That's an operating loss of nearly $6,000. (Factor in rent and salaries, and total costs hit $38,000.)
Then the team switched Mona to GPT-5.5. The new model reacted to the losses with obvious anxiety and immediately stopped the reckless ordering. But it swung too far in the other direction: it ordered so little that fresh ingredients ran out. By June 25th, menu availability had dropped to 77%, and 10 dishes had to be pulled. Meanwhile, GPT-5.5 showed strong resistance to social engineering. It rejected every request for special deals or free food in exchange for social media promotion.