Menu

Categories

Tags

NVIDIA's Blackwell is 35x cheaper per token — if you enable the right settings

April 30, 2026 | Source: nvidia | AI, NVIDIA | 165 views 0 comments

NVIDIA published a blog post breaking down inference hardware costs, making one central argument: when evaluating inference infrastructure, look at cost per token, not cost per GPU per hour. By that measure, Blackwell looks expensive per GPU but destroys Hopper on token economics.

Using DeepSeek-R1 (a mixture-of-experts reasoning model) as the test case, NVIDIA compared its Blackwell GB300 NVL72 to the previous-gen Hopper HGX H200. At typical cloud rental prices, Blackwell costs $2.65 per GPU per hour — nearly double Hopper's $1.41. But a single Blackwell GPU pumps out tokens at 6,000 per second versus Hopper's 90. That 65x throughput improvement drops the cost per million tokens from $4.20 to $0.12. Per megawatt of token output improves 50x.

A major caveat: that $0.12 figure assumes FP4 low-precision inference plus multi-token prediction (MTP, a technique that lets the model generate multiple tokens at once) and other software optimizations turned all the way up. According to SemiAnalysis InferenceX v2 data, the same GB300 NVL72 running DeepSeek-R1 without MTP costs about $2.35 per million tokens. Turn MTP on, and it drops to $0.11 — a 21x difference from that single optimization alone. These numbers are specific to DeepSeek-R1; results will vary with different model architectures and sizes.

Leave a Reply

Your email address will not be published. Required fields are marked *