Google DeepMind senior product manager and Google AI Studio product lead Logan Kilpatrick said on X that every company building products based on AI should create their own benchmarks (standardized test sets for measuring AI model performance). He calls this the way to make model improvements "disproportionately benefit your company" and advises founders and business owners to "start tomorrow."
https://twitter.com/OfficialLoganK/status/2048554074107470305
Currently, most companies choose AI models based on public leaderboards, but these leaderboards measure general capabilities and often disconnect from specific business scenarios. For example, a company doing contract review cares most about clause extraction accuracy, but public benchmarks don't test for that. As the DeepSeek V4 model recently acknowledged, even frontier models can have reasoning gaps that public benchmarks miss. So there's no way to know how good a model is at your specific task. The benefits of building your own benchmarks: first, each time a model is updated, you evaluate it with your own business tasks, selecting the best model for your scenario, not the one with the highest public ranking; second, you can feed these test sets back to model providers, pushing them to continuously optimize in the directions you care about.
Kilpatrick says Zapier, Sierra, and other companies are already doing this, and says "there's a lot of alpha to be created."