Menu

Categories

Tags

Cursor and Opus 4.7 edge out Codex in first AI programmer benchmark

May 12, 2026 | Source: artificialanalysis | AI, Developer | 293 views 0 comments

Artificial Analysis has released the first comprehensive benchmark index for coding agents, dubbed the Coding Agent Index. It combines three tests — code generation (SWE-Bench-Pro-Hard-AA), terminal operations (Terminal-Bench v2), and technical Q&A (SWE-Atlas-QnA) — to evaluate AI programmers' real-world engineering performance.

In the inaugural ranking, Cursor CLI paired with the Opus 4.7 model took the top spot with a score of 61, beating OpenAI's Codex (with GPT-5.5) and Anthropic's Claude Code (also with Opus 4.7) by a single point.

When both used the same Opus 4.7 model, Cursor CLI scored 61 versus Claude Code's 60 — but it came at a cost: average task time was longer (7.8 minutes vs. 5.8 minutes), and API calls were pricier ($1.47 vs. $1.24).

The cheapest option was Cursor's built-in Composer 2, at just $0.07 per task. DeepSeek V4 Pro ($0.35) and Kimi K2.6 ($0.76) followed.

But those domestic Chinese models took significantly longer. While Claude Code (with Opus 4.7) finished a single test task in as little as 5.8 minutes, DeepSeek V4 Pro averaged 18 minutes, and Kimi K2.6 dragged on to 41.5 minutes.

Leave a Reply

Your email address will not be published. Required fields are marked *