Vercel just dropped its June 2026 AI Gateway production index, and the numbers are wild. DeepSeek's token traffic share skyrocketed from less than 1% to 17% in a single month — overtaking OpenAI's 13% to claim third place. The catch? All that usage cost about 1% of the gateway's total spending.
It's all about pricing. DeepSeek V4 Flash charges just $0.14 per million input tokens and $0.28 per million output tokens. That's 20 to 50 times cheaper than Anthropic's comparable frontier models, and 8 to 12 times cheaper than rivals like Qwen 3.6 Plus and Kimi K2.6. Reviews show performance is solid, so dev teams are deploying it in production fast.
But cheap models don't necessarily mean cheap bills. Frontier models still dominate cash burn: Anthropic's spending share climbed from 61% to 65% in May, accounting for 70% to 80% of spending in high-stakes tasks like application generation, background agents, and coding. In coding agents, DeepSeek served 49% of tokens but only 4% of costs, while Anthropic handled 28% of tokens and gobbled up 70% of the money.
Teams are getting smart about it. They use intelligent routing to steer high-frequency, low-risk tasks to low-cost models and save frontier models for critical work. ROI math is also slowing upgrades. Google launched Gemini 3.5 Flash in May at a higher price than 3.0, and migration has been sluggish. At month's end, 3.0 Flash still carried 90% of Flash traffic; 3.5 Flash had just 7%. Meanwhile, AI agents are token hogs: they make up a quarter of requests but consume more than half of all tokens.