Menu

Categories

Tags

Capital One's BinEval framework uses yes/no questions to grade AI models

July 1, 2026 | Source: arxiv | AI | 209 views 0 comments

Researchers at Capital One have developed BinEval, an evaluation framework that automatically breaks down complex scoring criteria into specific yes/no questions, addressing the problem of black-box scoring and inflated grades. The framework has an evaluation model answer each yes/no question one by one, then calculates a score based on the proportion of correct answers.

In tests across three major datasets, BinEval using large language models like Claude Sonnet 4 matched or outperformed mainstream evaluation tools like UniEval, particularly good at catching answers that sound fluent but contain factual errors.

For example, in evaluating a summary about an aircraft interception, the summary read smoothly and got entities and aircraft models correct, but it swapped the positions of the Pentagon and Russia's accounts and fabricated a URL. An older AI judge, only looking at surface quality, gave a perfect 5.0. BinEval, with seven yes/no questions, accurately caught four factual errors and gave a score of 1.57 — very close to the human-assigned 2.0.

The error log from the yes/no questions can be used both to refine the judge model's own evaluation criteria and to automatically modify writing prompts. In instruction-following tests, feedback optimization improved compliance with format and sentence structure by 17 percentage points. However, the tool remains powerless against hard constraints like word count (which require mathematical calculation), and overly decomposing requirements can make evaluation standards too strict.

Tags: #Claude

Leave a Reply

Your email address will not be published. Required fields are marked *