
Scale AI's new benchmark, HarnessOpt-Bench, tests a very specific skill: whether an AI can modify another agent's harness and make it genuinely better.
A harness is the wrapper around the model — the prompts, tools, memory, and workflow that shape how the agent actually behaves. For the test, the AI gets a deliberately under-optimized harness. It looks at benchmark scores and failure cases, then iterates on the harness code. The model underneath never changes.
After the edits, Scale runs the agent through a set of test questions it has never seen before. That's how the company avoids an AI just gaming familiar benchmarks.
Scale ran 111 attempts across five models — Claude Opus 5, GPT-5.6 Sol, Kimi K3, and two others — across four task categories. Opus 5 came out on top, taking first place in three of the four.
The bigger takeaway: which model does the harness-rewriting matters a lot more than which coding tool it uses, whether that's Claude Code, Codex, or OpenCode.
Still, this isn't true AI self-improvement. It's one AI tweaking another agent's harness, and the starting harness was deliberately left with plenty of room to optimize.
Scale AI shared the results in a post on X.