
Zhipu AI's open-source model, GLM-5.2, has landed on the DeepSWE benchmark — a grueling test for long-horizon software engineering tasks — and it's already the top open-source entry. With its maximum thinking mode, the model achieves a 44% one-shot success rate on complex development tasks, 13 percentage points higher than the previous open-source leader, Kimi K2.7 Code.
GLM-5.2's average cost per task is $3.92, slightly above Kimi K2.7 Code's $2.82, but its success rate beats several major closed-source models in specific thinking configurations: Claude Sonnet 4.6 [high] (30%), Gemini 3.5 Flash [medium] (37%), and Claude Opus 4.8 [low] (41%).
Designed by Datacurve, the DeepSWE benchmark tests AI agents on long-horizon tasks with 113 real-world programming problems across five languages. Unlike traditional tests that modify a single spot of code, DeepSWE requires agents to make coordinated edits across multiple files, with an average fix touching over 600 lines of code. All tests run in isolated containers with strict CPU and memory limits.