NVIDIA RTX Spark Superchip Unveiled at Computex — Windows-on-Arm With 1 PetaFLOP and 128 GB Unified Memory
Jensen Huang revealed the RTX Spark at Computex 2026: a Windows-on-Arm superchip combining 20 Grace CPU cores, a Blackwell GPU, and up to 128 GB unified memory at 1 petaFLOP FP4. Laptops from Dell, HP, Lenovo, and Microsoft Surface arrive this fall starting at $1,499.
Jensen Huang took Computex on June 1 to announce the NVIDIA RTX Spark Superchip: a single chip combining 20 Grace ARM CPU cores, a Blackwell-generation GPU with 6,144 CUDA/RTX cores, and up to 128 GB of LPDDR5X unified memory at roughly 270–300 GB/s bandwidth. AI compute peaks at 1 petaFLOP FP4. GPU performance is comparable to an RTX 5070 laptop GPU. This is NVIDIA’s answer to Apple Silicon — and it runs Windows.
The positioning is explicit: CUDA comes to Windows on Arm. Every piece of NVIDIA’s AI software stack — TensorRT, cuDNN, DLSS, the full inference pipeline — now runs natively on Arm instruction sets without translation layers. That’s the part that matters to developers. Not just the hardware specs, but the ecosystem portability. Python ML workflows, local model inference, agent runtimes — if it runs on CUDA, it runs on RTX Spark.
Microsoft is a launch partner. The Surface Laptop Ultra ships with RTX Spark this fall. So do 30+ laptop models and 10+ desktop configurations across Dell, HP, Lenovo, ASUS, and MSI. Starting prices: Lenovo Yoga Slim 9i Spark Edition at $1,499, HP Spectre x360 16 at $1,699, Dell XPS 15 Arm Edition at $1,899. All fall 2026.
The 128 GB unified memory pool is the real differentiator against conventional discrete GPU laptop setups. Running a 70B parameter model in 4-bit quantization requires roughly 35–40 GB of contiguous memory. No laptop GPU currently ships with that. RTX Spark changes the math: you can run Llama 3.3 70B, a local Claude distill, or a Mistral 8x22B locally without offloading layers to CPU RAM. That’s not a benchmark number — that’s a capability that changes what offline, private, on-device AI actually means.
The tradeoff against Apple M5 Max is real: Apple Silicon’s memory bandwidth and power efficiency at the top tier are still advantages. But NVIDIA wins on software ecosystem depth — the CUDA library is 15+ years of GPU-optimized ML code, and porting that to Apple Metal remains incomplete for most research workloads.
NVIDIA described this as an “agentic AI OS” platform in partnership with Microsoft. That framing — not just AI hardware, but the chip as the substrate for autonomous agent workflows on Windows — is the marketing layer. The hardware specs justify the claim. One petaFLOP and 128 GB of fast unified memory is genuinely enough to run capable multi-step agentic pipelines locally.
Developer tooling and drivers launch alongside the first hardware units this fall.