HighPoint’s Rocket 7638D PCIe 5.0 switch card exploits Nvidia's GPUDirect interconnection to communicate with storage, bypassing CPU
Networking specialist HighPoint has launched its Rocket 7638D PCIe 5.0 switch card designed to enable direct interconnection between an Nvidia AI GPUs and NVMe storage devices, which could greatly reduce system lag when processing massive data sets.
Right now, for data to be processed in the GPU requires retrieval from the storage device and processing by the CPU, because only the CPU can talk to storage. The GPU cannot.
That is until Nvidia introduced the GPUDirect function in the Ampere generation of GPUs. This feature eliminates the need for a CPU and memory to act as the middleman and allows direct transfer of data from storage to the GPU. This serves to reduce the workload on the CPU and system memory and reduce latency in getting data to the GPU for processing.
[ Related: GPUs: Explaining the processing power behind AI ]
GPUDirect uses PCI Express Gen 5 and requires a PCIe switch that supports P2P DMA capability, and not all PCIe Gen5 switches support this feature. The Rocket card uses the Broadcom PEX 89048 switch, which system integrators to build systems with GPUDirect support.
The adapter features 48 PCIe 5.0 lanes, 16 of which are dedicated to internal NVMe storage devices while the rest are dedicated to connectivity. The adapter uses MCIO 8i connectors, which support up to 16 NVMe drives, for a total of up to 2PB of high-performance storage. It supports multiple GPU nodes by ensuring each GPU has dedicated bandwidth without sacrificing NVMe performance.
The Rocker 7638D adapter enables GPUDirect Storage workflows that avoid host CPU and RAM entirely and provide predictable bandwidth (up to 64 GB/s) and latency when paired with compatible software, which includes operating system, GPU drivers, and filesystem. The device is particularly useful in scenarios involving large-scale training datasets that use plenty of storage.
The adapter works out of the box with all major operating systems, without the need for special drivers or additional software installation. The Broadcom controller is based on an Arm design and is entirely self-contained and automated.
The adapter is targeted primarily at ML workloads. For model training, larger datasets stream from NVMe storage directly to GPUs without delays, reducing total training time. It can facilitate real-time data ingestion and fast GPU response, as well as enabling high-speed data augmentation and preparation.




