
DeepSeek has officially launched V4-Pro, and the leaked agent benchmark results are now confirmed. The final version is live across the app, web, and API. The company confirmed the scores that slipped out earlier: Terminal Bench 2.1 comes in at 87.9, DeepSWE jumped from 12.8 in the Preview to 62.7, and CyberGym hit 83.3. The release also adds native Responses API support, along with optimizations for Codex. Thinking intensity is now split into three tiers — low, high, and max — so users can dial it up or down based on task difficulty.