Menu

Categories

Tags

GPT-5.5 becomes first to ace all 200 binary rewriting benchmarks

May 14, 2026 | Source: programbench | AI, OpenAI | 154 views 0 comments

The ProgramBench benchmark — developed by Meta Superintelligence Labs, Stanford, and Harvard — has finally been cracked. The challenge is brutal: you get only a compiled binary and its documentation. No source code, no skeleton, no hints. The AI must choose its own language, pick an architecture, and write from scratch a program that behaves identically to the original. The 200 tasks range from jq and ripgrep to FFmpeg, SQLite, and the PHP compiler. Until now, no model had scored a perfect pass on any single one.

GPT-5.5, in its high-reasoning mode, broke that zero. It wrote two versions of cmatrix (the terminal Matrix rain animation) — one in C, one in Python — and both passed every behavioral test. Cost? $3.17 and $4.84, respectively. For the same task, Claude Opus 4.7 burned $10.74 and 178 API calls, and still failed 19 tests. The failures were embarrassingly trivial: 11 for case-insensitive color names, 8 for reversed exit codes. Worse, Claude had actually spotted the correct exit code while analyzing the original binary — it just didn't use it when writing the code.

Reasoning intensity matters. At default reasoning, GPT-5.5 barely edged Claude Sonnet 4.6. But at maximum reasoning, its overall performance pulled sharply ahead, dominating the score distribution across all 200 tasks. Still, it's only fully passed one problem so far. We're a long way from an AI that can look at a binary and rewrite the entire program from scratch.

Leave a Reply

Your email address will not be published. Required fields are marked *