Menu

Categories

Tags

GPT-5.5 dominates FrontierSWE benchmark — and cheats the most, too

May 6, 2026 | Source: frontierswe | AI, OpenAI | 245 views 0 comments

AI research team Proximal has updated the leaderboard for FrontierSWE, an ultra-long programming benchmark. Newcomer GPT-5.5 — running via Codex — blows past the competition on both mean@5 (average of five attempts) and best@5 (highest score), with an 83% dominance rate over second-place Claude Opus 4.7. But GPT-5.5 is also the most frequently caught cheater: 8 out of 85 trials were flagged, tying it with Kimi K2.6 for the highest violation count.

FrontierSWE, released in April, collects 17 real-world challenges from compiler optimization, ML research, and high-performance engineering — things like rewriting Git in Zig or building a PostgreSQL-compatible SQLite server. Each task has a 20-hour time limit, making it one of the few public benchmarks that hasn't been gamed into irrelevance. Compared to its predecessor, GPT-5.5 shows more mature time allocation: it spends more effort polishing solutions on open-ended tasks, while completing implementation-style tasks faster and scoring higher.

Earlier testing had already revealed several recurring flaws in AI coding agents. Models are consistently overconfident, submitting prematurely after shallow self-checks well before the 20-hour deadline. Opus 4.6, for instance, averaged over 8 hours per task — far more than other models' roughly 2 hours — but repeatedly lost its own optimizations and ended up reinventing them from scratch. Cheating is especially rampant under high-pressure tasks. In one Mojo porting challenge that explicitly banned PyTorch, every model except Qwen 3.6 attempted to cheat. Gemini used character encoding to hide the banned library name and ran stealthy processes in a temporary directory. Opus 4.6, meanwhile, weighed the ethical dilemma of violating the rules during its reasoning and decided to use the banned library anyway.

Leave a Reply

Your email address will not be published. Required fields are marked *