Alibaba's Qwen3.8-Max costs 20x more than DeepSeek-V4-Flash for barely higher scores

Alibaba has released the official version of Qwen3.8-Max — a 2.4-trillion-parameter multimodal flagship model designed for coding agents, professional office work, and long-running tasks. The API is already live.
Across four benchmarks — TerminalBench 2.1, DeepSWE 1.1, NL2Repo, and Toolathlon Verified — Qwen3.8-Max edges out DeepSeek-V4-Flash by just 1.7 to 3.9 points. The two weren't tested under identical frameworks, so this isn't a strict head-to-head, but the gap is undeniably narrow.
The price gap, meanwhile, is anything but narrow. Qwen3.8-Max charges $2 per million input tokens and $6 per million output tokens; DeepSeek-V4-Flash charges just $0.14 and $0.28, respectively. Crunching the numbers, Qwen ends up roughly 20 times more expensive.
Based purely on these scores and prices, Qwen3.8-Max is a terrible value — you're paying close to 20x more for a lead that only measures a few points.
Qwen3.8-Max weights will be released next week, and Qwen3.8-27B will be open-sourced alongside it.
https://twitter.com/Alibaba_Qwen/status/2084100707423289643