Large-scale AI workloads, with their high-bandwidth and low-latency requirements, are reshaping data center network architecture, Nvidia says.
Nvidia’s networking roadmap includes technology upgrades designed to transform data center compute, connectivity, and storage infrastructure into giant GPU-driven and network-defined AI factories.
Data centers are evolving into a new unit of computing, from a focus on CPUs to GPUs as the primary computing units and from the distribution of functions across different components to support the infrastructure for AI workloads, Gilad Shainer, senior vice president of networking at Nvidia, told Network World. This infrastructure evolution requires synchronized data movement and encompasses at least four networks: compute, scale-up, scale-out, and access.
[ Related: More Nvidia news and insights ]
“Today, the scale of the data center has shifted. In the era of AI, the data center itself has become the unit of computing. Instead of asking ‘How many CPUs can I buy?’ the question is now, ‘How do I design a data center capable of running my workloads at maximum efficiency?’ Shainer said. “This shift fundamentally changes how we design, connect, and optimize infrastructure. The giant data center has become the new unit of computing, and one current network architecture won’t handle,” he said.
“What’s needed is a layered design with bleeding-edge technologies — like co-packaged optics that once seemed like science fiction,” Shainer explained further in a recent blog.
“AI factories are rising — massive new data centers built not to serve up web pages or email, but to train and deploy intelligence itself,” Shainer wrote. “Welcome to the age of AI factories — where the rules are being rewritten and the wiring doesn’t look anything like the old internet. These aren’t typical hyperscale data centers. They’re something else entirely. Think of them as high-performance engines stitched together from tens to hundreds of thousands of GPUs — not just built, but orchestrated, operated and activated as a single unit. And that orchestration? It’s the whole game.”
It is this new data center network infrastructure that Nvidia has serious plans to develop now and in the future. Evidence of its intentions can be found in the vendor’s most recent announcement of extensions to Ethernet networking that will let widely distributed GPUs – in servers across multiple data centers, regardless of the distance between them – operate as a single, unified AI supercomputer.
Nvidia is baking into its Spectrum-X Ethernet platform a suite of algorithms that can implement networking protocols to allow Spectrum-X switches, ConnectX-8 SuperNICs, and systems with Blackwell GPUs to connect over wider distances without requiring hardware changes. These Spectrum-XGS algorithms use real-time telemetry — tracking traffic patterns, latency, congestion levels, and inter-site distances—to adjust controls dynamically.
Ethernet and InfiniBand
Developing and building Ethernet technology is a key part of Nvidia’s roadmap. Since it first introduced Spectrum-X in 2023, the vendor has rapidly made Ethernet a core development effort. This is in addition to InfiniBand development, which is still Nvidia’s bread-and-butter connectivity offering.
[ Related: What is an AI-optimized data center? ]
“InfiniBand was designed from the ground up for synchronous, high-performance computing — with features like RDMA to bypass CPU jitter, adaptive routing, and congestion control,” Shainer said. “It’s the gold standard for AI training at scale, connecting more than 270 of the world’s top supercomputers. Ethernet is catching up, but traditional Ethernet designs — built for telco, enterprise, or hyperscale cloud — aren’t optimized for AI’s unique demands,” Shainer said.
Most industry analysts predict Ethernet deployment for AI networking in enterprise and hyperscale deployments will increase in the next year; that makes Ethernet advancements a core direction for Nvidia and any vendor looking to offer AI connectivity options to customers.
“When we first initiated our coverage of AI back-end Networks in late 2023, the market was dominated by InfiniBand, holding over 80% share,” wrote Sameh Boujelbene, vice president of Dell ’Oro Group, in a recent report. “Despite its dominance, we have consistently predicted that Ethernet would ultimately prevail at scale. What is notable, however, is the rapid pace at which Ethernet gained ground in AI back-end networks. As the industry moves to 800 Gbps and beyond, we believe Ethernet is now firmly positioned to overtake InfiniBand in these high-performance deployments.”
In the coming year or two, “Ethernet will become the more prominent technology for networking, surpassing InfiniBand as 800G ramps and 1.6T take form,” predicts 650 Group in a recent report. “The 800G cycle for AI will set records for revenue and ports.”
“Spectrum‑X is fully standards‑based Ethernet. In addition to supporting Cumulus Linux, it supports the open‑source SONiC network operating system — giving customers flexibility,” Shainer wrote. “A key ingredient is Nvidia SuperNICs — based on Nvidia BlueField-3 or ConnectX-8 — which provide up to 800 Gb/s RoCE connectivity and offload packet reordering and congestion management.”
“Spectrum-X brings InfiniBand’s best innovations — like telemetry-driven congestion control, adaptive load balancing and direct data placement — to Ethernet, enabling enterprises to scale to hundreds of thousands of GPUs,” Shainer wrote. “Large-scale systems with Spectrum‑X, including the world’s most colossal AI supercomputer, have achieved 95% data throughput with zero application latency degradation. Standard Ethernet fabrics would deliver only ~60% throughput due to flow collisions.”
Copper and optics
Optics is an important element in the connectivity for scale up networks, for Nvidia’s NVLink, because of the huge amount of bandwidth that runs between those connected GPU silicon devices.
“We are focusing on increasing the compute density inside of a single rack, so that single rack could use copper. Copper consumes zero power, it’s reliable, very cost effective. As long as you can use copper, use copper. But when customers scale out to greater distances, copper cannot be used because it cannot run those distances, and this is where need to go optics,” Shainer said.
NVLink is the company’s flagship high-speed interconnect technology, born out of its Mellanox networking group, which lets multiple GPUs in a system or rack share compute and memory resources, thus making many GPUs appear to the system as a single processor.
Currently, NVLink features up to 1.8 TB/s of bi-directional bandwidth per GPU, supporting up to 72 GPUs per rack. Faster and higher capacity NVLink technology is expected to evolve rapidly in the new few years to handle higher speeds and even more GPU-to-GPU communications.
In the optical arena, Nvidia offers pluggable optics for its Ethernet and InfiniBand networking gear. But this summer, the vendor said it would jump headlong into the co-packaged optics (CPO) networking world. CPO integrates network optic components directly into a switch ASIC. CPO technology is expected to rapidly develop in the coming months and years to handle AI traffic and ultimately other network traffic that demands high performance.
For its part, Nvidia’s CPO-based products will come in the form of Spectrum-X Photonics for Ethernet and Quantum-X InfiniBand Photonics.
“Nvidia has designed CPO-based systems to meet unprecedented AI factory demands. By integrating optical engines directly onto the switch ASIC, the new Nvidia Quantum-X Photonics and Spectrum-X Photonics will replace legacy pluggable transceivers,” wrote Ashkan Seyedi, director of product marketing, in a blog about the new systems in August. “The new offerings streamline the signal path for enhanced performance, efficiency, and reliability. These innovations not only set new records in bandwidth and port density but also fundamentally alter the economics and physical design of AI data centers.”
The result of Nvidia’s CPO suite is a 3.5x leap in power efficiency compared to previous architectures, and a 10x improvement in resiliency by reducing the number of overall optical components that may fail, Seyedi stated.
Nvidia’s networking business
Ultimately, workloads are continuing to advance, and they require more computing capabilities and bigger computing infrastructure, Shainer said.
“The AI factory has become a single unit, and the infrastructure, or the network, defines what that single unit can do. So that makes the infrastructure a critical item, and it needs to be designed in a consistent way, it needs to be designed together with everything else. This is where you will see innovation coming, innovation in software framework, especially in hardware science. You will see network functions being run on the NIC and GPU and more,” Shainer said.
The importance of networking for Nvidia’s business is already showing up in its financial reporting.
The Futurum Group noted after Nvidia’s most recent quarterly financial call that the vendor’s networking was a record bright spot, with revenue almost doubling year-over-year to $7.3 billion, driven by strong adoption of Spectrum-X Ethernet, InfiniBand XDR, and NVLink scale-up systems.
“Spectrum-X surpassed a $10 billion annualized run rate, and NVIDIA unveiled Spectrum-XGS, designed to interconnect multiple giga-scale AI factories. InfiniBand revenue nearly doubled sequentially, reflecting its role in low-latency networking for leading model makers. NVLink, now in its fifth generation, provides 14x the bandwidth of PCIe Gen 5, a critical enabler of rack-scale Blackwell systems. With networking efficiency gains directly translating to significant customer savings, this business is becoming a critical margin driver and reinforces Nvidia’s positioning as a full-stack AI infrastructure vendor,” Futurum wrote.
Networking emerged as the most impressive growth driver in Q3, Futurum wrote. “As of today, we still believe this segment remains underappreciated and underestimated by many. We expect investors’ awareness of the importance of networking, including scale-up, scale-out, and scale across, as well as the comparative advantage it provides to Nvida’s overall edge over its competitors and attractiveness to customers,” Futurum stated.




