Menu

Categories

Tags

Claude Opus 5 takes the crown as AI agents start tuning each other's harnesses

August 19, 2026 | Source: t | AI, Developer | 113 views 0 comments

Scale AI's new benchmark, HarnessOpt-Bench, tests a very specific skill: whether an AI can modify another agent's harness and make it genuinely better.

A harness is the wrapper around the model — the prompts, tools, memory, and workflow that shape how the agent actually behaves. For the test, the AI gets a deliberately under-optimized harness. It looks at benchmark scores and failure cases, then iterates on the harness code. The model underneath never changes.

After the edits, Scale runs the agent through a set of test questions it has never seen before. That's how the company avoids an AI just gaming familiar benchmarks.

Scale ran 111 attempts across five models — Claude Opus 5, GPT-5.6 Sol, Kimi K3, and two others — across four task categories. Opus 5 came out on top, taking first place in three of the four.

The bigger takeaway: which model does the harness-rewriting matters a lot more than which coding tool it uses, whether that's Claude Code, Codex, or OpenCode.

Still, this isn't true AI self-improvement. It's one AI tweaking another agent's harness, and the starting harness was deliberately left with plenty of room to optimize.

Scale AI shared the results in a post on X.

Leave a Reply

Your email address will not be published. Required fields are marked *