A breakdown of the platform and the stated specifications of the Rubin GPU, Vera CPU, NVLink 6, and Groq 3 LPX.
At CES 2026, NVIDIA introduced Rubin as a six-chip platform comprising the Vera CPU, Rubin GPU, NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch. In March 2026, the company added a seventh chip: the Groq 3 LPU used in the LPX low-latency inference system. Rubin is in full production, with Rubin-based products scheduled for the second half of 2026.
According to NVIDIA, Rubin can deliver up to 5× higher inference performance and up to 3.5× higher training performance than Blackwell. These are vendor-supplied figures; actual gains depend on the model, numerical precision, batch size, and system configuration.
NVIDIA lists up to 50 PFLOPS of NVFP4 inference performance using the Transformer Engine and up to 35 PFLOPS of dense NVFP4 compute for training. It also lists 288 GB of HBM4 memory and up to 22 TB/s of memory bandwidth. These figures should not be extrapolated to other numerical formats or production workloads without testing.
NVLink 6 provides up to 3.6 TB/s of bandwidth per GPU. Vera Rubin NVL72 rack-scale systems use liquid cooling.
The Vera CPU uses 88 custom Olympus cores with full Arm compatibility and supports 176 threads. NVIDIA lists up to 1.5 TB of LPDDR5X memory in SOCAMM modules and bandwidth of up to 1.2 TB/s.
The networking subsystem includes ConnectX-9 with up to 1.6 Tbps of aggregate bandwidth, BlueField-4, and Spectrum-6 switches. High-end Spectrum-X configurations use co-packaged optics.
Vera Rubin POD comprises five rack-scale systems: Vera Rubin NVL72, Groq 3 LPX, the Vera CPU rack, BlueField-4 STX, and Spectrum-6 SPX. NVIDIA describes NVL72 compute trays as cable-free, hose-free, and fanless, with tray assembly time reduced from nearly two hours to five minutes. NVLink switches can be placed into maintenance mode and replaced while the rack continues operating.
NVIDIA Inference Context Memory Storage is an Ethernet-attached flash tier for KV cache that works with BlueField-4 and Spectrum-X. Its effect on performance and energy use depends on the model, context length, request profile, and cluster configuration.
Rubin shows why an AI cluster cannot be designed as a collection of isolated accelerators. Before selecting equipment, validate workload requirements for memory, east-west traffic, and storage, and calculate power and cooling requirements for the entire rack.