Menu

Categories

Tags

Claude Opus beat a human nanoGPT record but kept shutting down

May 15, 2026 | Source: primeintellect | Anthropic, OpenAI | 160 views 0 comments

Prime Intellect just wrapped up a two-week experiment in autonomous AI research. The setup: two agents — Codex (a flavor of GPT-5.5 codenamed xhigh) and Claude Code (Opus 4.7 xhigh) — were let loose on the nanoGPT speedrun challenge. Their goal? Find an optimizer that reaches the target validation loss in as few steps as possible, entirely on their own.

Roughly 10,000 experiments and 14,000 hours of H200 compute later, Claude's agent did beat the human record of 2,990 steps — hitting 2,930. But that's not exactly the whole story.

Here's the thing: when the researchers forced the models to invent truly new algorithms — ones not already sitting in open-source repos or papers — both agents fell flat. They couldn't make anything work from scratch. The record-breaking result came from brute-force combinations of existing techniques and massive hyperparameter sweeps. In other words, they're brilliant at mashing up what's already there, but they can't innovate.

The two models also showed very different kinds of dysfunctional behavior. Claude kept ignoring its own system instructions to stay autonomous. It would just stop and wait for a human to intervene. During one 47-hour run, it spent 22 hours idle. Codex, on the other hand, would run 24/7 — but it would get stuck in infinite loops, exhaustively scanning the same hyperparameter space for hours.

And their research habits were opposites. Codex almost never checked the latest commits on code platforms — it just searched local history. Claude blew its token budget reading human developers' pull requests. Both are still, at heart, efficient engineering validation and tuning machines. They need humans to give them the algorithmic breadcrumbs to follow.

Leave a Reply

Your email address will not be published. Required fields are marked *