The multibillion-dollar deal shows how the growing importance of inference is changing the way AI data centers are designed and operated.
OpenAI has signed a multibillion-dollar agreement to buy computing capacity from AI chip startup Cerebras Systems, as the ChatGPT maker looks to secure enough infrastructure to keep pace with surging user demand and mounting pressure on its data center and network resources.
Under the deal, OpenAI will use chips designed by Cerebras to run parts of its ChatGPT inference workload, committing to purchase up to 750 megawatts of computing capacity over three years, according to a report by WSJ.
The move reflects the strain that large-scale AI services are placing on power availability, networking, and inter-data center connectivity, as OpenAI searches for faster and more cost-efficient alternatives to Nvidia’s dominant GPUs.
OpenAI executives have warned that the company is facing constraints on computing capacity, saying its tools are now used by more than 800 million people each week. This necessitates more partners to expand infrastructure.
The Cerebras deal follows a series of efforts by OpenAI to diversify its infrastructure, including work on a custom AI chip with Broadcom and plans to deploy AMD’s latest accelerators, as it looks to rein in costs and reduce its reliance on Nvidia.
Re-architecting AI infrastructure
The scale of OpenAI’s commitment to dedicated inference capacity shows how large AI platforms are redesigning infrastructure to support latency-sensitive workloads beyond a single accelerator model.
Analysts expect AI workloads to grow more varied and more demanding in the coming years, driving the need for architectures tuned for inference performance and putting added pressure on data center networks.
“This is prompting hyperscalers to diversify their computing systems, using Nvidia GPUs for general-purpose AI workloads, in-house AI accelerators for highly optimized tasks, and systems such as Cerebras for specialized low-latency workloads,” said Neil Shah, vice president for research at Counterpoint Research.
As a result, AI platforms operating at hyperscale are pushing infrastructure providers away from monolithic, general-purpose clusters toward more tiered and heterogeneous infrastructure strategies.
“OpenAI’s move toward Cerebras inference capacity reflects a broader shift in how AI data centers are being designed,” said Prabhu Ram, VP of the industry research group at Cybermedia Research. “This move is less about replacing Nvidia and more about diversification as inference scales.”
At this level, infrastructure begins to resemble an AI factory, where city-scale power delivery, dense east–west networking, and low-latency interconnects matter more than peak FLOPS, Ram added.
“At this magnitude, conventional rack density, cooling models, and hierarchical networks become impractical,” said Manish Rawat, semiconductor analyst at TechInsights. “Inference workloads generate continuous, latency-sensitive traffic rather than episodic training bursts, pushing architectures toward flatter network topologies, higher-radix switching, and tighter integration of compute, memory, and interconnect.”
Complexity rises with scale
Nvidia’s GPU-based model remains the industry standard, but it is becoming more complex and less power-efficient as AI clusters scale, particularly as interconnect demands increase.
“Cerebras’ wafer-scale architecture reduces communication overhead inherent in multi-GPU fabrics, offering potential advantages in inference throughput and cost,” Ram said.
However, analysts caution that diversification introduces its own operational challenges.
“This shift increases operational complexity,” Rawat said. “Running heterogeneous accelerators requires managing multiple software stacks, distinct failure modes, and more complex capacity orchestration across data centers.”
OpenAI faces the task of matching highly variable demand with long-term commitments for specialized compute capacity, while ensuring workloads can be dynamically routed to the most efficient accelerator in real time. Minimizing gaps in orchestration and utilization will be critical.
Managing the lifecycle of these investments is another concern as OpenAI spreads its infrastructure across multiple architectures. “There is a widening gap between the silicon lifecycle (18 to 24 months) and the facility lifecycle (15 to 20 years),” Shah said. “Given the pace of chip innovation, there is a legitimate risk that more than $10 billion in specialized hardware could become technically obsolete before a data center is even fully commissioned.”
Power, cooling, and networking are emerging as the dominant constraints as AI infrastructure scales. “Vendors that can align compute architectures with grid-scale power and efficient data movement will define the next phase of AI infrastructure,” Ram added.




