Menu

Categories

Tags

AI's true abilities are being underestimated by 60%, new report warns

July 4, 2026 | Source: aisi | AI | 474 views 0 comments

A new report from the UK AI Safety Institute argues that mainstream AI agent testing has a major blind spot: by capping the compute budget during evaluations, researchers are dramatically underestimating how capable and how quickly these models are evolving.

The team tested several leading large models on benchmarks in cybersecurity, software engineering, and mathematics. Their findings show that an agent's performance isn't a fixed score — it's a curve that keeps climbing as you throw more test-time compute at it. For example, in cyber attack/defense tasks, when the compute budget was raised from 2.5 million tokens to 50 million tokens, the most advanced agents could tackle tasks equivalent to 14 hours of human work — up from just 2 hours at the lower limit. Many attempts that failed with limited compute succeeded once the agents were given enough runway to explore and correct mistakes.

Newer models also use that test-time compute far more efficiently than older ones. When evaluated with ample budgets, the measured capability evolution trend (the slope of the fitted curve) is about 60% steeper than under low-compute testing, suggesting traditional evaluations seriously understate AI's real iteration speed. However, this compute dividend has limits: in fields like medicine where immediate feedback is scarce, more compute doesn't boost agent performance.

As inference costs drop, relying on low-budget benchmarks could lead policymakers to underestimate the risks posed by AI agents in real-world deployments.

Leave a Reply

Your email address will not be published. Required fields are marked *