At CES, Nvidia launches Vera Rubin platform for AI data centers

News
Jan 6, 20265 mins

The Nvidia Rubin platform consists of six new chips, including four networking processors.

Portrait of Two Diverse Developers Working on Computers, Typing Lines of Code that Appear on Big Screens Surrounding Them. Male and Female Programmers Creating Innovative Software, Fixing Bugs.
Credit: Gorodenkoff / Shutterstock

Nvidia used the Consumer Electronics Show (CES) as the backdrop for an enterprise-scale announcement: It launched the Vera Rubin NVL72 server rack platform for AI data centers, featuring new concepts and technology like “context memory” storage, zero downtime maintenance, rack-scale confidential computing, and several other advancements.

As with previous naming conventions, the Vera Rubin platform is named after a notable scientist, in this case, astronomer Vera Rubin, who proved the theory of dark matter. Vera is the codename for an Arm-based high-performance CPU and Rubin is the name of the GPU platform.

But one thing Nvidia CEO Jensen Huang stressed in his keynote is that the Vera Rubin platform is actually 6 pieces of silicon. In addition to the CPU and GPU, there are four networking processors: NVLink 6 Switch, ConnectX-9 SuperNIC, BlueField-4 DPU, and Spectrum-6 Ethernet Switch.

Each Vera CPU comes with 88 custom cores, 176 threads using spatial multi-threading, 1.5 TB LPDDR5x system memory, and 1.2 Tbps memory bandwidth, alongside confidential computing and a 1.8 Tbps NVLink C2C interconnect. That’s seven times the bandwidth of PCI Gen 6, said Dion Harris, Nvidia’s senior director of HPC and AI infrastructure in a pre-briefing ahead of the show.

“Vera doubles data processing, compression and code compilation performance versus our prior gen Grace CPU across both training and inference,” he said. “These Vera features maximize GPU utilization by orchestrating, routing, scheduling the KV cache and contacts.”

The Vera CPU is the first to support confidential computing to deliver the Trusted Execution Environment, which helps maintain data security across CPU, GPU and the MV link domain, protecting the world’s largest proprietary models, training data and inference workloads.

Rubin GPUs can deliver 50 petaflops for inference using NVFP4 data format—five times faster than Blackwell—and hit 35 petaflops for NVFP4 training, which is 3.5 times faster than Blackwell. HBM4 memory offers 22 Tbps bandwidth — 2.8x over Blackwell — and NVLink bandwidth per GPU is 3.6 Tbps, double Blackwell’s speed.

Networking is enhanced with the liquid-cooled NVLink 6 Switch, offering 400G SerDes, 3.6 Tbps per-GPU bandwidth, total switching bandwidth of 28.8 Tbps, and 14.4 teraflops of FP8 in-network compute capability.

The complete platform gives the Vera Rubin NVL72 platform up to 3.6 exaflops of NVFP4 inference, which is five times faster than the previous generation platform, and up to 2.5 exaflops of NVFP4 training, 3.5 times higher than the previous generation.

Vera Rubin NVL72 includes 54 TB of LPDDR5x capacity (2.5x Blackwell), 20.7 TB HBM4 (50% more), 1.6 Pbps HBM4 bandwidth (2.8x increase), and a scale-up bandwidth of 260 Tbps (double that of Blackwell NVL72). “That’s more bandwidth than the entire global Internet,” said Harris.

Nvidia also redesigned the rack, announcing its Third-Gen NVL72 Rack Resiliency. Features include a cable-free modular tray design that enables assembly and servicing 18 times faster than the previous generation.

The NVLink Intelligent Resiliency feature supports server maintenance with “zero downtime,” keeping racks operational even during component swaps or partial population. The second-generation RAS Engine allows for GPU diagnostics without taking the rack offline.

Initially, Rubin will be offered in two formats: the Vera Rubin NVL72 rack-scale platform. featuring 72 Rubin GPUs and 36 Vera CPUs, and the HGX Rubin NVL8 platform with eight Rubin GPUs for use with x86-based servers.

Not surprisingly, Nvidia announce that a wide swath of leading tech firms plan to support Rubin, Including Amazon Web Services (AWS), Anthropic, Cisco, Cohere, CoreWeave, Dell Technologies, Google, HPE, Lenovo, Meta, Microsoft, Nebius, Nscale, OpenAI, Oracle Cloud Infrastructure (OCI), Perplexity, Supermicro, and xAI.

To handle the massive data sets generated by agentic AI,  Nvidia is introducing a new storage platform it said will offer a significant boost in inference performance and power efficiency, called the Nvidia Inference Context Memory Storage Platform.

The platform uses BlueField-4 and Spectrum-X Ethernet tightly coupled with the Nvidia Dynamo and Nixl, enabling coordinated context retrieval across memory, storage and networking.

“The result is a big step forward in performance and efficiency compared to traditional network storage used for inference context,” said Harris. This platform delivers up to 5x higher tokens per second, 5x better performance per TCO dollar and 5x better power efficiency than traditional storage systems.

“That translates directly into higher throughput, lower latency and more predictable behavior. And it really matters for the workloads we’ve been talking about, large context applications like multiturn chat retrieval, augmented generation and a and agentic AI, multistep reasoning,” said Harris.

Nvidia said it is collaborating with its storage partners to integrate this inference context memory into the Rubin platform, aiming to offer customers a comprehensive, unified AI infrastructure.

Andy Patrizio is a freelance journalist based in southern California who has covered the computer industry for 20 years and has built every x86 PC he’s ever owned, laptops not included.

Andy writes the Data Center Explorer blog for Network World. His work has appeared in a variety of publications, including Tom's Guide, Wired, Dr. Dobbs Journal, Tech Target, Business Insider, and Data Center Knowledge. Earlier in his career, he held editorial positions at IT publications like InternetNews, PC Week and InformationWeek.

Andy holds a BA in Journalism from the University of Rhode Island.

More from this author