
Meta AI, Stanford, and Harvard researchers have released ProgramBench, a new coding benchmark from the same team behind SWE-bench. The task: take a compiled binary and its documentation, then reconstruct a complete codebase from scratch that replicates the program's behavior — no source code, no decompilation, no internet access, and no restrictions on language or architecture.
The benchmark includes 200 tasks ranging from small CLI tools like jq and ripgrep to massive projects like FFmpeg, SQLite, and the PHP interpreter. In total, there are over 248,000 behavioral tests, automatically generated by agent-driven fuzz testing. The evaluation uses mini-SWE-agent as a unified baseline without per-task harness tuning.
Results? Every single one of the 9 frontier models tested failed completely. On the primary metric — passing every test case — none managed a single success. On the secondary metric of near-passing (95% or more tests passed), Claude Opus 4.7 led with 3%, followed by Claude Opus 4.6 at 2.5%. The rest — Claude Sonnet 4.6, Claude Haiku 4.5, Gemini 3.1 Pro, Gemini 3 Flash, GPT 5.4, GPT 5.4 mini, and GPT 5 mini — all scored 0%. The authors note that most agents didn't time out; they declared themselves done and submitted implementations that were behaviorally incomplete. ProgramBench code is open source under MIT license.