Back to Blog
AI Models March 30, 2026 5 min read

ARC-AGI-3 Offers $2M to Any AI That Can Think Like a Human — Every Frontier Model Scores Below 1%

The ARC Prize Foundation just released its hardest benchmark yet: an interactive reasoning test where GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro all score under 1%, while untrained humans score 100%. The $2M prize remains unclaimed.

ARC-AGI-3 Offers $2M to Any AI That Can Think Like a Human — Every Frontier Model Scores Below 1%

The ARC Prize Foundation released ARC-AGI-3 on March 25, and the results are a cold bucket of water on the AGI narrative. Every frontier model tested — GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, Grok-4.20 — scored below 1%. Humans, with no training, score 100%. The prize pool exceeds $2 million. It’s still sitting there.

This is not a failure of measurement. It is a measurement of failure.

ARC-AGI-3 is a radical departure from its predecessors. The original benchmark used static grid puzzles to test abstract pattern recognition. Version 3 abandons grids entirely and drops AI agents into interactive, video-game-like environments. There are no instructions. The agent must observe its environment, infer the rules, set its own goals, and solve problems it has never seen before. François Chollet, who created ARC and leads the foundation, describes this as testing fluid intelligence — the capacity to adapt reasoning to genuinely novel situations.

The specific scores: Gemini 3.1 Pro Preview hit 0.37%, GPT-5.4 reached 0.26%, Claude Opus 4.6 managed 0.25%, and Grok-4.20 scored exactly 0.00%. These are models that routinely score above 90% on bar exams, coding benchmarks, and graduate-level science problems. On ARC-AGI-3, they are functionally indistinguishable from noise.

What makes this especially pointed is that simple convolutional neural networks and graph-search approaches reach 12.58% on the same benchmark. Systems with far fewer parameters, built for far narrower purposes, outperform the most capable language models in the world on this specific task. That is not a margin of error. It is a structural signal about what scaling transformer architectures actually learns.

The prize structure: $700,000 for the first agent to score 100% on the full evaluation set; additional tracks for top performers and efficient solutions. No time limit has been set. At current trajectories, it may be a while.

The benchmark’s design deliberately resists the optimization games that have corrupted earlier evals. Because the environments are procedurally generated and interactive, there is no training set to overfit. An agent that scores well here cannot have memorized the answers. It has to actually think — or produce something that functions like thinking.

The AI industry has a habit of moving goalposts. When ARC-AGI-1 was solved, the response was to say that benchmarks don’t measure what matters. ARC-AGI-3 makes that argument harder. Untrained humans, given no hints, no fine-tuning, no chain-of-thought prompting, solve every environment. The most powerful AI systems built by humanity cannot solve 1% of them.

That gap is the actual frontier. Everything else is extrapolation.

ARC-AGI benchmarks AGI reasoning