Cerebras Unveils the CS-4: Three Wafer-Scale Chips, 750 Petaflops, and a Claimed 30x Speed Advantage Over GPUs
The CS-4 packs three WSE-3 Turbo wafers — 4 trillion transistors and 900,000 cores each — into one system with 129.6 PB/s of memory bandwidth. Cerebras claims inference up to 30x faster than GPU systems at roughly half the rack power of Nvidia and AMD.
Cerebras announced the CS-4, its first multi-wafer system: three Wafer Scale Engine 3 Turbo (WSE-3T) processors working as one machine, delivering a combined 750 petaflops of sparse FP16 compute and a claimed inference speed “up to 30 times faster than GPU-based systems” on frontier models.
The per-wafer numbers remain absurd by conventional chip standards. Each WSE-3T carries 4 trillion transistors and 900,000 AI-optimized cores with 44GB of on-chip SRAM, fabbed on TSMC’s 5nm process. Stitched together, the three wafers expose 129.6 petabytes per second of memory bandwidth, and wafer-to-wafer latency drops from five microseconds to two. Cerebras says the system supports models exceeding 50 trillion parameters. In its benchmark, the CS-4 pushed gpt-oss-120b past 4,400 tokens per second per user — the kind of decode speed that turns an agentic loop from minutes into seconds.
Power is the quieter headline. Cerebras estimates 120–140 kilowatts per rack, roughly half of comparable Nvidia and AMD systems, underpinning its claim of a tenfold throughput-per-watt advantage. In a market where data center buildouts are increasingly constrained by electricity rather than silicon supply, performance per watt is becoming the metric that closes deals.
CEO Andrew Feldman framed the pitch around latency: “In AI, speed is productivity.” CTO Sean Lie was more specific about why it matters now — the speed boost “gives an order of magnitude more reasoning, verification, or tool use.” That is the agentic-AI argument in one sentence: when a coding agent makes hundreds of sequential model calls, per-token latency compounds into the difference between an interactive tool and a batch job.
The honest caveat: the WSE-3T is not new silicon. It is the existing WSE-3 with its clock roughly doubled, from about 1.4GHz to 2.8GHz — an aggressive binning-and-cooling exercise rather than a redesigned chip, as critics were quick to note. The genuinely new generation is slated for 2027. And Cerebras disclosed no pricing and no customer commitments at launch, though shipping began in Q3 2026.
Still, the timing is pointed. Nvidia’s grip on AI hardware is strongest in training; inference — where workloads are exploding as agents go into production — is where wafer-scale architecture has always had the cleanest story. Doubling throughput without waiting for new silicon, at half the power draw, is exactly the move a challenger makes when the incumbent’s customers are power-constrained and latency-sensitive. Whether the 30x claim survives independent benchmarks is the number to watch.