
Security firm Aikido put seven models through their paces with a code-auditing agent, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, and GPT-5.6 Sol, Luna, and Terra. The test used 32 recently disclosed vulnerabilities, and each model got three runs. Qwen3.8-Max found 26 of those bugs across its three attempts, a recall rate of 81.3% that ties Claude Opus 5 for first place. Its F1 score, which balances false negatives and false positives, landed at 83.2% — slightly below …