SemiAnalysis: AI tokens cost $0.99/million and account for 30% of salaries
Agentic AI is wreaking havoc on the unit economics of professional services. Research firm SemiAnalysis reveals that internal large language model token spending now consumes 30% of total employee salaries. The average worker blasts through nearly 5 billion tokens per month; top contributors burn over 100 billion. Tasks that used to take analysts hours — like converting Excel models or building financial charts — now get done in minutes for a few bucks in token costs.
https://twitter.com/SemiAnalysis_/status/2070915305858007345
The secret sauce is the effective cost collapse. Sure, Opus 4.7 lists for $5 per million input tokens and $25 per million output. But agent workflows boast a ridiculous 300:1 input-to-output ratio, and over 90% prompt cache hit rates. The blended rate? Just $0.99 per million tokens.
Software and hardware are doubling down. Running DeepSeek R1 on B300 with wideEP, disagg, and MTP optimizations pushes single-GPU throughput from 1,000 to 14,000 tokens per second — a 14x software boost. On hardware, an optimally configured GB300 NVL72 churns out 17 times the throughput of an H100 (32x with FP4). That’s a structural cushion for LLM developer margins, and a clear signal that token prices by 2027 will make today look expensive.