Google Unveils 8th-Gen TPU 8t and 8i at Cloud Next — Built for the Agentic AI Era
At Google Cloud Next 2026, Google announced two purpose-built eighth-generation TPUs: TPU 8t for training at 9,600-chip superpod scale, and TPU 8i for low-latency agent inference. Both chips were co-designed with Google DeepMind.
Google split its TPU line in two. At Google Cloud Next 2026 in Las Vegas, the company announced the eighth generation of its Tensor Processing Units as two distinct chips built for diverging workloads: TPU 8t for training, and TPU 8i for inference. Both were co-designed with Google DeepMind and are purpose-built for the agentic AI era.
That split is the headline. Every previous TPU generation was a single chip architecture asked to do both jobs. The 8th gen acknowledges that training and inference have fundamentally different requirements and that a chip optimal for one is a compromise on the other.
TPU 8t — training at superpod scale
TPU 8t scales to 9,600 chips in a single superpod, sharing 2 petabytes of high-bandwidth memory across the pod. Google claims nearly 3x the compute performance of Ironwood (the previous gen) and 2x better performance-per-watt. The stated impact: frontier model development cycles shrinking from months to weeks. That’s the number that matters for Google’s own AI research velocity and for cloud customers training large models.
TPU 8i — inference for concurrent agents
TPU 8i is built for throughput and latency, not raw training compute. A single pod connects 1,152 chips with 3x more on-chip SRAM than Ironwood. The SRAM density matters for KV cache sizes — the bigger the cache, the longer the context window you can serve without memory pressure. Google is explicitly targeting workloads that need to run millions of concurrent agents. With the agent layer of AI products exploding in 2026, that’s the right bet.
The architectural split signals something important about where Google thinks AI infrastructure is going. Training is becoming a more concentrated, scheduled workload — fewer massive runs, more orchestrated pipelines. Inference, by contrast, is becoming spikier, more latency-sensitive, and increasingly multi-modal. One chip can’t serve both cleanly.
On competitive positioning: NVIDIA’s Blackwell Ultra handles both training and inference with the same die, leaning on scale and interconnect for flexibility. Google’s two-chip approach is a deliberate differentiation — claiming that specialization beats generalism at the frontier.
Cloud Next’s other announcements — Workspace Intelligence, expanded Gemini 2.5 integrations, and GKE AI inference optimization — were significant, but the TPU 8t/8i split is the strategic move. Google is building its own silicon moat against a future where NVIDIA’s pricing power is the biggest line item on every AI company’s P&L.
Availability wasn’t announced at the keynote. Expect a preview program to start with existing Google Cloud customers running large-scale AI workloads.
Related reading
- Hardware Google Signs Marvell as Its Third Custom AI Chip Partner to Co-Design MPU and Inference TPU
- Hardware Cerebras Files for $26.6B Nasdaq IPO — The Wafer-Scale AI Chip That's 57x Bigger Than Nvidia's H100
- Hardware Amazon's Chip Business Hits $20B Run Rate — Jassy Says It Could Be the Next Nvidia, and Plans to Sell Externally