Menu

Categories

Tags

AI benchmarks are dead, argues OpenAI's Noam Brown — inference compute matters most

June 10, 2026 | Source: x | AI, OpenAI | 219 views 0 comments

OpenAI researcher Noam Brown has a message for the AI evaluation community: static benchmark scores are becoming meaningless for frontier models. Instead, he argues, we need to look at inference compute scaling curves — plots of performance against the amount of compute used during inference.

Brown uses the hypothetical model GPT-5.5 as an example. In standard tests, GPT-5.5 barely outperforms GPT-5.4. But give it more inference compute, and its performance explodes. Standard tests with limited compute budgets simply can't capture a model's true ceiling.

This pattern is backed up by external evaluations. In Andrej Karpathy's autonomous agent research and in cybersecurity tests by the UK AI Safety Institute, both GPT-5.5 and the Mythos model kept improving as inference budgets increased — even after generating over 100 million tokens, with no ceiling in sight. The more powerful the model, the more it benefits from additional compute.

All this compute also complicates safety evaluations. Brown warns that current biological and cybersecurity benchmarks typically don't fix an inference budget. If a state-level adversary throws $10 million of compute at a specific task, a model that seems safe in standard tests could cross dangerous red lines.

Brown's recommendation: when releasing new models, AI companies should publish performance curves with tokens, compute cost, or runtime on the x-axis. Evaluators need to treat inference budget as a core variable, and in safety assessments, use low-compute tests to extrapolate safety boundaries.

Leave a Reply

Your email address will not be published. Required fields are marked *