xAI Sets Public Launch of Grok 4.5, a 1.5-Trillion-Parameter 'Opus-Class' Model
Elon Musk confirmed Grok 4.5, built on xAI's new V9 foundation model, launches publicly this week after private beta testing at SpaceX and Tesla. A 2-trillion-parameter follow-up is already targeting August.
Elon Musk confirmed that Grok 4.5 launches publicly this week, following private beta testing across SpaceX and Tesla. It’s built on xAI’s new V9 foundation model at 1.5 trillion parameters — a 50% scale increase over Grok 4.4, which shipped in late May.
What’s under the hood
Musk is calling Grok 4.5 “Opus-class,” directly invoking Anthropic’s flagship as the comparison point, and claims it delivers faster inference, greater token efficiency, and lower operating cost than that benchmark. xAI hasn’t published third-party benchmark numbers alongside the announcement, so that positioning is a claim to verify once independent evals land, not a confirmed result.
The predecessor, Grok 4.3, shipped in April; Grok 4.4 followed in late May. Musk has now signaled the roadmap doesn’t stop at 4.5 — a 2-trillion-parameter model is already targeting an August 2026 launch, and xAI reportedly has seven models in training simultaneously as it works toward a rumored 10-trillion-parameter Grok 5.
The SpaceXAI context
Grok 4.5’s rollout lands right after xAI’s rebrand signal to SpaceXAI on July 6, folding the AI company’s public identity closer to Musk’s satellite and rocket business. The user-facing products stay the same — Grok, Grok Build, Imagine, Voice, and Grokipedia — but the launch cadence suggests xAI is leaning on that combined infrastructure and capital base to fund back-to-back trillion-parameter training runs at a pace none of its competitors are currently matching.
Why it matters
A 50% parameter increase in roughly six weeks is an aggressive release cycle by any frontier-lab standard — Anthropic and OpenAI typically space flagship updates months apart. Whether that pace produces genuine capability gains or diminishing returns per parameter is the open question; scaling alone stopped being a reliable proxy for quality once every lab hit compute and data limits over the past two years. For developers evaluating model choice, the practical takeaway is to wait for independent benchmarks — Terminal-Bench, SWE-Bench, GPQA — before shifting production workloads, since self-reported “Opus-class” claims from any lab have historically needed discounting.