Nvidia Launches Vera CPU: 88 Cores Built to Win the Agentic AI Bottleneck
Nvidia's new Vera CPU bets on single-threaded performance, not raw core count, claiming 1.8x the throughput of x86 rivals on agentic workloads. A successor codenamed Rosa is already on the roadmap.
Nvidia doesn’t think the agentic AI bottleneck is the GPU. On July 7, the company launched Vera, an 88-core CPU it’s calling “the max single-threaded CPU at scale” — a deliberate break from the industry’s core-count arms race.
The logic: agentic systems spend most of their critical path on the CPU, not the GPU. Tool calling, code execution, data processing, KV-cache management, and result analysis all run on general-purpose cores while the accelerator waits. If that work is serial and latency-bound, adding more cores doesn’t help — you need each core to finish its job faster. Vera is built on a custom architecture Nvidia calls the Olympus core, engineered to maximize per-core performance rather than chase core count.
The numbers: 88 cores with SMT support for 176 threads, 1.2 TB/s of bandwidth to LPDDR5X memory, and a monolithic compute die delivering 3.4 TB/s of core-to-core bandwidth to keep those cores fed. Nvidia claims 1.8x higher throughput than x86 competitors on “loaded CPU workloads that represent agentic execution,” 1.5x on coding workflows, and 3x faster database analytics — at twice the power efficiency and 50% faster than traditional CPU designs, according to the company.
Nvidia also confirmed the next generation is already scoped: a CPU codenamed Rosa, built on a new Rigel core, sits on the roadmap behind Vera. That’s a signal this isn’t a one-off SKU — Nvidia is committing to CPU design as a standing product line, not a bolt-on to its GPU business.
The timing matters. This lands the same week CNBC reported Nvidia’s Kyber rack system for Rubin Ultra has slipped to 2028 over manufacturing issues — a rare public stumble that opened room for AMD and Google to make ground. Vera is Nvidia reasserting where it thinks the real differentiation is: not the accelerator alone, but the full agentic pipeline, CPU included. Pairing Vera with its GPU roadmap under the “Vera Rubin” platform name makes the CPU a first-class part of the pitch to hyperscalers building agent infrastructure, not an afterthought sourced from Intel or AMD.
For teams building agent infrastructure, the practical takeaway is that CPU selection is about to become a real performance lever again, not a commodity decision. If your agent stack is bottlenecked on tool-call latency or KV-cache thrashing rather than raw model FLOPs, Vera-class hardware is aimed squarely at that problem — expect early availability through Nvidia’s DSX AI factory partners before broader enterprise rollout.