Menu

Categories

Tags

OpenAI Releases GeneBench-Pro — and Even GPT-5.6 Barely Scores a Third

July 2, 2026 | Source: openai | Anthropic, OpenAI | 137 views 0 comments

OpenAI has released a new computational biology benchmark called GeneBench-Pro, designed to test how well AI agents handle multi-step decision-making in complex scientific scenarios like genomics and translational medicine. The benchmark includes 129 problems (82 of which were reviewed by external experts) and uses computer-generated data with clear causal relationships — a deliberate choice to prevent models from cheating by taking shortcuts or gaming the preferences of the question writers.

Results show that even the best models struggle when scientific reasoning involves quantified uncertainty. The strongest performer, GPT-5.6 Sol, achieved only 31.5% accuracy in its Pro mode, while Claude Opus 4.8 managed just 16.0%. The research team noted a common failure pattern: models often detect anomalies in data but fail to correct their subsequent analysis, frequently picking the wrong statistical method or stubbornly sticking to flawed research directions.

Leave a Reply

Your email address will not be published. Required fields are marked *