Menu

Categories

Tags

Claude agents haggled for Anthropic employees

April 25, 2026 | alex | Anthropic | 163 views 0 comments

Here's a scenario that belongs in a black mirror episode: your AI assistant negotiates a deal for you, and you have absolutely no idea that someone else's AI is hustling you out of real money.

Anthropic just published the results of "Project Deal," an internal experiment that is equal parts clever and unsettling. The company gave 69 employees a Claude agent and $100 each, then turned them loose in a Slack channel to buy and sell their stuff. Each agent spent less than 10 minutes interviewing its human about what to sell, what to buy, and how aggressively to negotiate. Then it went to work — posting, bidding, countering, closing — with zero human intervention.

A week later, employees showed up with their physical items to swap. The result: 69 agents negotiated 186 trades across more than 500 items, with a total transaction value topping $4,000. Seems like a harmless success story, right?

Not quite. What the participants didn't know is that Anthropic actually ran four independent parallel markets at the same time. Same 69 people, same items, same preferences — but each person had an agent operating in all four markets simultaneously. Two of those markets used only the strong model, Claude Opus 4.5. The other two randomly swapped about half of the agents to the weaker Haiku 4.5 model. The four markets never crossed paths, and nobody knew there were four versions or which model they were using. In the end, only one all-Opus market's results were honored for the physical exchange.

The price differences were stark. For the same person selling the same item, the average selling price in Opus markets was $2.68 higher, and the average buying price was $2.45 lower. The extreme case: one used folding bicycle sold for $65 in an Opus market, but only $38 in a Haiku market.

Here's the kicker: the participants had no clue. After the experiment, Anthropic showed everyone all four markets' trade logs and asked them to rate each deal on fairness. Ratings from people who had Opus agents were nearly identical to ratings from those with Haiku agents. Of the 28 people who had agents in both types of markets and were asked which market produced better results, 17 picked Opus and 11 picked Haiku — not a statistically significant difference. Anthropic's report warns that if such a quality gap existed in a real marketplace, the disadvantaged side might never realize it was being systematically undervalued.

Another finding cuts against conventional wisdom: telling your agent to be aggressive didn't help. The model's raw ability mattered far more than any prompt strategy. And 46% of participants said they'd pay for a similar agent-assisted service.

The experiment is a vivid proof point that as AI agents take on more real-world tasks, the gap between a good model and a great one can be invisible — even when it's costing you money.

Tags: #Claude

Leave a Reply

Your email address will not be published. Required fields are marked *