Menu

Categories

Tags

AI Can't Be an Autonomous Scientist Yet — CUSP Benchmark Shows It Lacks Foresight

June 11, 2026 | Source: t | AI, Anthropic | 267 views 0 comments

Sakana AI, Stanford University, Oxford University, the Allen Institute for AI, and the Alan Turing Institute have jointly introduced a temporal benchmark called CUSP, designed to evaluate AI's ability to predict scientific progress. The benchmark includes 4,760 scientific milestones from journals like Nature and Science, which are broken down into 17,429 specific tasks. It tests frontier large language models — including GPT-5.4, Claude Sonnet 4.5, and DeepSeek R1 — on their ability to forecast unknown scientific futures.

Large language models typically answer questions by rote memorization of their training data. So directly testing them on whether they can predict truly novel discoveries requires a carefully designed benchmark like CUSP. The results so far suggest that AI still cannot function as an autonomous scientist, lacking the forward-looking research vision that humans possess.

Leave a Reply

Your email address will not be published. Required fields are marked *