Infrinia AI Cloud OS automates Kubernetes and inference services on GPU infrastructure.
SoftBank has launched Infrinia AI Cloud OS, a software stack for operating AI data centers that automates infrastructure management and provides inference services for large language models.
The software handles tasks from BIOS configuration to Kubernetes management on GPU platforms, including Nvidia’s GB200 NVL72.
“By deploying Infrinia AI Cloud OS, AI data center operators can build Kubernetes as a Service (KaaS) in a multi-tenant environment, and Inference as a Service (Inf-aaS) that provides Large Language Model inference capabilities via APIs, as part of their own GPU cloud services,” SoftBank said in a statement.
The company said it developed the software to address the operational complexity involved in running GPU cloud services.
In addition, the software stack is expected to reduce total cost of ownership (TCO) as well as operational burden compared with bespoke solutions or in-house development, the company added in the statement.
The launch marks SoftBank’s expansion beyond hardware into the GPU cloud software layer, according to Charlie Dai, VP and principal analyst at Forrester. “This elevates SoftBank from pure infrastructure operator to AI-native platform-level competitor,” Dai said.
Addressing enterprise challenges
The software provides two main services, according to SoftBank. The Kubernetes-as-a-Service component automates the stack from BIOS and RAID settings through the OS, GPU drivers, networking, Kubernetes controllers, and storage, the company said.
It reconfigures physical connectivity using Nvidia NVLink and memory allocation as users create, update, or delete clusters, according to the announcement. The system allocates nodes based on GPU proximity and NVLink domain configuration to reduce latency, SoftBank said.
Enterprises currently face complex GPU cluster provisioning, Kubernetes lifecycle management, inference scaling, and infrastructure tuning challenges that require deep expertise, according to Dai.
SoftBank’s automated approach addresses these pain points by handling BIOS-to-Kubernetes configuration, optimizing GPU interconnects, and abstracting inference into API-based services, he said. This allows teams to focus on model development rather than infrastructure maintenance, Dai said.
The Inference-as-a-Service component lets users deploy inference services by selecting large language models without configuring Kubernetes or underlying infrastructure, according to the company. It provides OpenAI-compatible APIs and scales across multiple nodes on platforms including the GB200 NVL72, SoftBank said.
The software includes tenant isolation through encrypted communications, automated system monitoring and failover, and APIs for connecting to portal, customer management, and billing systems, according to the announcement.
Growing market, intensifying competition
The launch positions SoftBank to compete in a market projected to grow from $8.21 billion in 2025 to $26.62 billion by 2030.
SoftBank faces competition from hyperscale cloud providers and specialized GPU vendors. AWS, Microsoft Azure, and Google Cloud offer managed Kubernetes services with GPU support through EKS, AKS, and GKE, respectively. Specialized providers, including CoreWeave, Lambda Labs, and RunPod, have built Kubernetes-native platforms targeting similar operational challenges.
CoreWeave operates 45,000 GPUs and is Nvidia’s first Elite-level cloud services provider. Lambda Labs generated $425 million in revenue in 2024 and offers H100 instances at $2.49 per hour, according to Contrary Research.
SoftBank’s software-centric approach signals a shift in competitive advantage from GPU availability to platform automation, according to Dai. “As GPU-as-a-Service demand accelerates, differentiation increasingly depends on intelligent orchestration, inference abstraction, and integrated AI lifecycle tooling,” he said. The market is moving toward full-stack AI-native cloud platforms rather than raw compute provisioning, Dai said.
Rollout strategy
SoftBank plans to first deploy the software in its own GPU cloud services before expanding to external customers. The Infrinia team aims to deploy the software to overseas data centers and cloud environments, the company said.
“The advancement of AI infrastructure requires not only physical components such as GPU servers and storage, but also software that integrates these resources and enables them to be delivered flexibly and at scale,” Junichi Miyakawa, SoftBank’s president and CEO, said in a statement. SoftBank said the software aims to reduce the total cost of ownership and operational burden compared with custom solutions or in-house development. The company did not disclose pricing or availability details.




