Agentic AI is reshaping the unit economics of professional services. Research firm SemiAnalysis reports that internal LLM token spending now accounts for 30% of total employee compensation, with the average user consuming nearly 5 billion tokens per month — and top contributors burning through over 100 billion. Tasks that once took analysts hours, like converting Excel models or generating earnings charts, now take minutes at a cost of just a few dollars in tokens.
https://twitter.com/SemiAnalysis_/status/2070915305858007345
The real game-changer is the plummeting actual cost per token. While Opus 4.7 lists at a steep $5 per million input tokens and $25 per million output, real-world agent tasks have an extreme input-to-output ratio of 300:1, and prompt caching hits above 90%. That brings the blended cost to just $0.99 per million tokens.
Software and hardware improvements are accelerating the cost collapse. Running DeepSeek R1 on B300, software optimizations like wide expert parallelism (wideEP), disaggregated serving (disagg), and multi-token prediction (MTP) boost single-GPU throughput from a baseline of 1,000 tokens/sec to 14,000 — a 14x software-only gain. On the hardware side, a fully optimized GB300 NVL72 delivers 17x the throughput of an H100 (32x with FP4), structurally protecting LLM developer margins and pointing to token prices far below today's by 2027.