In a joint interview with AWS CEO Matt Garman on Stratechery, OpenAI CEO Sam Altman made a provocative prediction: per-token pricing for AI models is on its way out.
Altman used GPT-5.5 as an example — the model costs more per token than GPT-5.4, but it uses far fewer tokens to complete the same task. And customers, he argued, don't care about token counts. They care about whether the job is done and what the total price is. He framed OpenAI not as a token factory but as an "intelligence factory," where clients want to buy as much intelligence as possible for the least money, regardless of the underlying model or how many tokens are consumed.
Altman also noted that far more customers are asking OpenAI for increased capacity than for price cuts. He drew a contrast with utilities: "Water is cheap, but you won't drink much more of it. With intelligence, if the price is low enough, I'll just keep using more — no other utility works like that." Garman chimed in, pointing out that the unit cost of compute has dropped by orders of magnitude over the past 30 years, yet more compute is being sold than ever.
When tech commentator Ben Thompson brought up Altman's 2023 prediction that AI would threaten Google Search, Altman offered a candid reassessment. The outcome was better than expected, he said, but in unexpected ways. He had assumed AI would transform how people find information online and directly disrupt search. Instead, OpenAI's success has come primarily from ChatGPT — which he called "the first truly mass-market consumer product since Facebook" — and the Codex API. Search itself has not been disrupted. He acknowledged that Google remains underestimated in many respects.
The interview, which focused on the newly announced Bedrock Managed Agents partnership between OpenAI and AWS, touched on broader themes of AI commoditization and the shift from raw compute to outcomes. OpenAI's recent model releases, including GPT-5.5, have already demonstrated the trend toward fewer tokens for the same output. Meanwhile, the API pricing for GPT-5.5 went live with a per-token premium that reflects the model's efficiency — a pricing structure that may soon feel as outdated as the dial-up modem.