Developers are using DeepSeek V4’s absurdly cheap API to run massive automated coding sessions — the kind of thing OpenAI is also targeting with its Codex browser control. The key? Extremely high cache hit rates that bring per-token costs down to almost nothing.
Users have posted screenshots of their bills. One developer ran a V4 Pro model for two and a half hours to auto-fix CI errors, consuming 80 million tokens. Thanks to a 99.41 percent cache hit rate, the total cost was just 4 RMB — about 55 cents. Another developer burned through 27.8 billion tokens in a single day and paid only $160. (That likely used the V4 Flash model, based on the pricing.) For comparison, doing the same with Claude Sonnet 4.6, assuming the same cache hit rate, would cost roughly $11,076. That’s a difference of over $10,900.
How does DeepSeek get away with this? Extreme price cuts. V4 Pro is currently on a limited-time 75 percent discount (extended to May 31, dropping output pricing to $0.87 per million tokens). The company also permanently slashed cache hit prices by 90 percent across all models: V4 Pro’s cache hit price is now $0.003625 per million tokens, and V4 Flash is even lower at $0.0028. In agentic coding scenarios where the same code prefix is loaded repeatedly, cache hit rates skyrocket. At those throughput levels, pure pay-as-you-go API pricing actually beats fixed monthly subscriptions with call limits.
To handle the flood of agent traffic, DeepSeek updated its third-party integration guides: set the model name to deepseek-v4-pro[1m] in Claude Code for a million-token context window, and OpenCode and OpenClaw now support it natively — much like Zhipu AI’s GLM-5V-Turbo which plugs into the same frameworks.