OpenAI just released GeneBench-Pro, a new benchmark designed to test AI agents on complex genomics and translational medicine tasks that require multi-step decision-making. The benchmark includes 129 problems (82 of which were reviewed by external experts), and uses computer-simulated data with clear causal relationships to prevent models from gaming the system by taking shortcuts or catering to the test creators' biases.
The results? Even the best models struggle with scientific reasoning that involves quantifiable uncertainty. OpenAI’s most powerful model, GPT-5.6 Sol, managed only about 30% accuracy in its full-capability mode.