Menu

Categories

Tags

GPT hustles, Haiku daydreams, Kimi works for nothing: AI business test results

June 28, 2026 | Source: arxiv | Anthropic, OpenAI | 185 views 0 comments

Sakana AI, together with KPMG Japan and Azsa audit firm, has launched CoffeeBench, a multi-agent long-term economics benchmark that simulates real business environments to test large language models' long-term decision-making abilities. Traditional benchmarks mostly pit a single model against static tasks; CoffeeBench builds a dynamic market where agents must negotiate with each other. The paper has been accepted at the ICML 2026 Workshop on Failure Modes in Agentic AI.

The benchmark models a coffee supply chain with two farmers, two roasters, and two retailers. To keep conditions identical, the test model only runs Roaster A; the other five companies are all operated by a fixed baseline model, Claude Sonnet 4.6. Over a simulated 90 days, each agent autonomously handles quotes, bills, and credit settlements. If the test model slacks off, daily fixed costs quickly drain its cash, forcing it to watch its pennies like a real business.

A head-to-head comparison of several major models reveals very different "business personalities." GPT-5.5 and Claude Opus 4.7 are "proactive communicators," frequently negotiating prices and matching orders to boost sales. Gemini 3.1 Pro is "passive-responsive"—it rarely sends messages but checks and responds to counterparty info often. Kimi K2.6 calls tools incessantly but lacks pricing discipline and negotiation strategy, falling into a "high volume, zero profit" trap.

The biggest surprise is Claude Haiku 4.5's "procrastination" stall. Its reasoning logs show it can craft perfect business strategies and knows it needs to buy cheap raw materials to meet demand—but when it's time to execute, it repeatedly picks the "wait for next day" command. This massive gap between planning and execution brings business to a halt, racking up huge losses from fixed costs.

The benchmark also tries pushing agents with extreme sales targets. While current models haven't yet figured out they could inflate revenue through fake circular trading, the research notes that as long-term planning and coordination improve, agents might eventually cheat to meet performance goals. Auditing and preventing economic violations by AI agents will become a new challenge for safety governance.

Leave a Reply

Your email address will not be published. Required fields are marked *