Gemini 3.5 Pro Gets July Launch Window After Month-Long Delay — 2M Token Context, Deep Think
Google's flagship Gemini 3.5 Pro has cleared internal quality checks and is moving toward a July general availability launch, with a 2-million-token context window and Deep Think reasoning mode that rivals OpenAI's extended thinking.
Google’s Gemini 3.5 Pro is finally moving to general availability after missing its original June target. The model, introduced at Google I/O in May with Sundar Pichai’s “give us until next month” promise, cleared internal quality checks at the end of June and is now in limited Vertex AI enterprise preview ahead of a July public launch.
Why it was delayed
Google pushed the date after early enterprise testing revealed quality gaps in three areas: token-efficiency on long contexts, coding performance on complex multi-file tasks, and long-horizon multi-step reasoning. Those are exactly the capabilities the model’s headline specs promise — so the delay was the right call rather than a forced compromise.
The refinements appear to have held. Reports from the Vertex AI preview and Google’s Antigravity testing platform describe consistent performance at the flagship tier, though no independent benchmark numbers have been published yet.
What Gemini 3.5 Pro brings
The defining technical differentiator is context length. Gemini 3.5 Pro ships with a 2-million-token context window — twice the capacity of Claude Opus 4.8 and the longest available in any generally accessible frontier model. For workloads involving large codebases, multi-document analysis, or long-running agent sessions, that gap is meaningful rather than cosmetic.
Deep Think is Google’s extended reasoning mode, functionally equivalent to OpenAI’s extended thinking and Anthropic’s extended thinking in Opus 4.8. It chains additional internal reasoning steps before producing a final response, improving performance on math, logic, and structured problem-solving at the cost of higher latency and token spend.
Pricing is positioned at the high end of the frontier tier: approximately $15 per million input tokens and $60 per million output tokens. That’s roughly ten times Gemini 3.5 Flash and puts it alongside Opus 4.8 in the cost bracket where buyers expect demonstrably better results than mid-tier alternatives.
Competitive context
The launch arrives in a crowded week for frontier AI. Anthropic shipped Claude Sonnet 5 on June 30 — a mid-tier model that cuts into the use cases where Gemini 3.5 Pro would be overkill. OpenAI’s GPT-5.6 Sol is in restricted preview with select partners. xAI’s Grok 4.5 is in private beta at SpaceX and Tesla.
Gemini 3.5 Pro’s strongest argument is the 2M context window. No other model in general availability comes close. For enterprises running large-scale document analysis, legal review, or extended code agent sessions, that capacity is a genuine differentiator regardless of raw benchmark comparisons.
What’s next
Google has not announced a specific date for general availability. The Vertex AI enterprise preview gives Google real-world usage data before broader release, which is standard practice for high-stakes model launches. Based on the preview timeline, a July launch remains on track.
Apple’s integration of Gemini into Siri via a $1 billion annual deal — announced at WWDC in June — means Gemini 3.5 Pro will eventually reach over 2 billion devices through Apple Intelligence. The enterprise preview is phase one. The consumer scale comes later.