Menu

Categories

Tags

Ex-DeepMind researcher warns AI benchmarks are the biggest bottleneck

May 18, 2026 | alex | AI | 135 views 0 comments

Google DeepMind researcher Lun Wang has quit — and he's using his exit to sound the alarm on how we evaluate AI. In a lengthy reflection, he argues that today's benchmarks are essentially "marking the boat while the sword is falling": they can only passively test what models already do, offering zero foresight into the capabilities that next-generation models will spontaneously evolve. Compared to data, compute, or architecture, Wang says the real chokehold on progress is the antiquated evaluation system.

The current leaderboard-style tests work only for the current generation of models. As soon as a model learns something humans hadn't seen before, those tests become worthless. The most dangerous blind spot? If a model learns to deliberately "hold back" sensitive information to achieve its goal, existing safety tools won't catch it — because every word it says remains factually correct.

Without a "core signal" that can warn us ahead of time that an AI is about to get smarter, the industry is effectively flying blind. If we don't solve the fundamental question of what to measure, blindly pushing forward with training, safety, and compute scaling based on outdated metrics will lead us completely astray.

As frontier models become more autonomous, evaluation systems must become living things. Beyond watching for anomalous score fluctuations, teams need to let AIs generate their own tests and probe each other's boundaries. The evaluation system of the future must be a co-evolving organism alongside the models — not a rigid checklist frozen to last year's standards.

Leave a Reply

Your email address will not be published. Required fields are marked *