Menu

Categories

Tags

ARC-AGI-3 benchmark splits into two leaderboards: raw AI scores below 1%, agents hit 36%

July 30, 2026 | Source: t | AI, OpenAI | 162 views 0 comments

The ARC-AGI-3 paper introduces two separate leaderboards. The official leaderboard bans external harnesses (agent frameworks), requiring all models to use the same minimal prompt with no tools — it's a test of raw, scaffold-free intelligence. The community leaderboard, on the other hand, allows harnesses, scores are self-reported, and the ARC Prize defaults to no verification. The paper explicitly warns that "scores on the community leaderboard should not be interpreted as evidence of AGI progress."

On the official leaderboard, frontier models all score below 1%: Gemini 3.1 Pro Preview achieves 0.37%, GPT...

Leave a Reply

Your email address will not be published. Required fields are marked *