Nvidia’s Rubin CPX chip combines Vera CPUs, Rubin GPUs and is aimed at massive-context processing
Nvidia has taken the wraps off a new purpose-built GPU along with a next-generation platform specifically targeted at massive-context processing as well as token software coding and generative video.
The Rubin CPX chip is a derivative of Nvidia’s next-generation Rubin GPU (the successors to the Blackwell GPU) for massive context inference. In terms of practical use, Rubin CPX is focused on the highest performance and token revenue for long-context processing.
Generative AI providers like ChatGPT, Google Gemini, and Perplexity sell their services through the sale of tokens, and their models do the processing. A simple query might cost 100 tokens while a complex reasoning query might cost over 100X more tokens. The faster and more efficiently provider can process tokens, the more revenue it generates.
Inference is often considered to be a single step in the AI process, but it’s two workloads, according to Shar Narasimhan, director of product in Nvidia’s Data Center group. They are the context or prefill phase and the decode phase. Each of these two phases has different requirements of the underlying AI infrastructure.
The context (or prefill) phase is compute-intensive, whereas the decode (or generation) phase is memory-intensive, but up to now, the GPU has been asked to do both when it really does only one task well. The Rubin CPX has been engineered to improve compute performance specific to the context phase, Narasimhan said.
“It will dramatically increase the productivity and performance of AI factories,” said Narasimhan. It achieves this through massive token generation. Tokens equal work units in AI, particularly generative AI, so the more tokens generated, the more revenue generated.
Rubin has two dies with 25 petaFLOPs per die, NVLink interconnect and 288GB of HBM4 high-speed memory. The Rubin CPX has one die with 30 petaFLOPS of performance, no NVLink and 128GB of GDDR7 memory. So Rubin CPX is optimal for specific high context needs that don’t need a lot of memory. CPX will be cheaper than the standard Rubin but Nvidia would not say how much.
To process video, AI models can take up to one million tokens for an hour of content, which can take many hours if not days to generate. The more tokens the system can generate, the larger scale processing it can do.
Rubin CPX delivers up to 30 petaflops of compute with NVFP4 precision. It features 128GB of GDDR7 memory rather than the usual HBM memory, which is more expensive than GDDR7. Nvidia says that the GDDR7 has adequate performance, and that Rubin CPX delivers three times faster attention capabilities compared with GB300 NVL72 systems.
Rubin CPX is offered in multiple configurations, including the Vera Rubin NVL144 CPX, that can be combined with the Quantum‐X800 InfiniBand scale-out compute fabric or the Spectrum-XTM Ethernet networking platform with Nvidia Spectrum-XGS Ethernet technology and Nvidia ConnectX-9 SuperNICs.
Nvidia is also announcing a new Vera Rubin NVL 144 CPX rack. Narasimhan said the NVL 144 CPX enables AI service providers to dramatically increase their profitability by delivering $5 billion of revenue for every $100 million invested in infrastructure.
It comes in two configurations: single rack, with 144 Rubin CPX GPUs, 144 Rubin GPUs, and 36 Vera CPUs for 8 exaFLOPs of NVFP4 compute and 100TB of fast memory and 1.7 PB/s of memory bandwidth. Nvidia said it is 7.5 times faster than the current top-of-the-line GB300 NVL72.
The other configuration is the dual-rack system where Vera CPUs and Rubin GPUs are in one rack augmented by a dedicated rack of Vera Rubin CPX to do context (prefill). So customers can purchase racks without CPX servers, with CPX servers mixed in, and CPX servers in a second, separate rack,
Nvidia Rubin CPX is expected to be available at the end of 2026.
More Nvidia news and insights:
- Nvidia networking roadmap: Ethernet, InfiniBand, co-packaged optics will shape data center of the future
- Nvidia’s new computer gives AI brains to robots
- Nvidia turns to software to speed up its data center networking hardware for AI
- Nvidia: ‘Graphics 3.0’ will drive physical AI productivity
- Nvidia launches Blackwell-powered RTX Pro GPUs for compact AI workstations




