Back to Blog
AI Models March 30, 2026 5 min read

Google Launches Gemini 3 Deep Think: 41% on Humanity's Last Exam, Live for Ultra Subscribers

Google's Gemini 3 Deep Think is now available to AI Ultra subscribers, hitting 41.0% on Humanity's Last Exam and 45.1% on ARC-AGI-2 — using parallel reasoning to explore multiple hypotheses simultaneously.

Google Launches Gemini 3 Deep Think: 41% on Humanity's Last Exam, Live for Ultra Subscribers

Google launched Gemini 3 Deep Think for AI Ultra subscribers today, posting 41.0% on Humanity’s Last Exam without tools and 45.1% on ARC-AGI-2 with code execution. Both are claimed as industry-leading scores on two of the hardest public reasoning benchmarks.

The model is live in the Gemini app now. To access it: select “Deep Think” in the prompt bar, then choose “Gemini 3 Pro” from the model dropdown. Early API access is also available for researchers and enterprise customers.

What makes Deep Think different

The key technical claim is “advanced parallel reasoning” — the model explores multiple hypotheses simultaneously rather than generating a single chain of thought. For complex math and science problems, this means the model can hold competing solution paths in parallel, evaluate them, and converge on the most defensible answer. That’s architecturally different from chain-of-thought prompting where the model commits to a reasoning path early.

Google describes this as designed specifically for complex math, science, and logic problems — not as a general-purpose upgrade. This scoping is honest and useful. Deep Think isn’t faster or cheaper; it’s designed for hard problems where standard reasoning fails.

The benchmark context

Humanity’s Last Exam is a collection of 3,000 expert-level questions across mathematics, physics, chemistry, biology, law, and other fields — designed to be unsolvable by current AI systems. A 41.0% score is notable because the test was built to be hard, not because 41% is a high number in absolute terms.

ARC-AGI-2 is more controversial as a benchmark — it tests reasoning on novel visual patterns designed to require genuine generalization rather than memorization. A 45.1% score with code execution is strong by current standards.

For comparison, Gemini 2.5 Deep Think achieved gold-medal performance at the International Mathematical Olympiad and ICPC World Finals. Gemini 3 extends that trajectory into broader scientific reasoning.

The Ultra subscriber strategy

Requiring an Ultra subscription for access is consistent with Google’s strategy of using leading models as conversion drivers for premium tiers. The question is whether Deep Think’s capabilities are distinctive enough to justify the subscription upgrade versus using OpenAI o3 or Anthropic Claude 3.7 Sonnet on comparable tasks.

For researchers doing formal mathematics or scientific modeling, the parallel reasoning architecture is a genuine capability differentiator. For most developers and knowledge workers, the practical difference may be smaller than the benchmark numbers suggest. Early API access for enterprise customers is the more interesting signal — it suggests Google is positioning Deep Think as infrastructure for high-stakes scientific and engineering workflows, not just a chat model upgrade.

Google Gemini AI Reasoning AI Models