Anirban Ghoshal
Senior Writer

Alibaba Cloud tweaks software for networking efficiency gains

News
Sep 2, 20254 mins

Researchers at the Chinese cloud giant have published details of their networking software optimizations.

Alibaba Cloud the front view
Credit: Alibaba Cloud

Alibaba Cloud has figured out ways to improve network throughput in its data centers without adding more hardware, and next week its researchers will present their discoveries  at the ACM SIGCOMM 2025 conference.

The network improvements they will present to the Association for Computing Machinery’s Special Interest Group on Data Communication have been tested in live networks over the last year or two, and include ZooRoute, a fast failure recovery service that Alibaba Cloud claims can ensure global bypass in large-scale cloud networks within seconds.

“When failures occur, strategies like fast reroute and traffic engineering can only offer local bypass in seconds or global bypass in minutes. Though protective rerouting and network architecture simplifications are proposed to accelerate global reconvergence, they require upgrades to underlying equipment, making large-scale deployment challenging,” company researchers wrote in a paper presenting  ZooRoute.

If data center operators are unwilling to make the necessary equipment upgrades, “tenants are forced to develop their own recovery solutions, which typically involve redundant resources or protocol stack modifications, thereby increasing capital and operating expenses,” the researchers wrote.

In contrast, ZooRoute can continuously probe for viable paths and reroute traffic instantly during failures, with no need for changes to hardware or to tenant applications, they said in the paper.

Charlie Dai, vice president and principal analyst at Forrester, said that where other cloud service providers (CSPs) such as AWS and Google use fast reroute and traffic engineering, ZooRoute’s proactive approach appears more aggressive and fine-grained than typical solutions.

Alibaba Cloud said that it has been using ZooRoute in AliCloud for the last 18 months, where it has reduced outage time by 92.71%.

Nezha for network performance in high-demand VMs

Another software upgrade is helping Alibaba Cloud maintain network performance for high-demand virtual machines (VMs) without spending more on SmartNIC-accelerated virtual switches (vSwitches).

Nezha, a distributed vSwitch load-sharing system, identifies idle SmartNICs and uses them to create a remote resource pool for high-demand virtual NICs (vNICs).

Alibaba has tested the system in its data centers for a year and said in the paper that “Nezha effectively resolves vSwitch overloads and removes it as a bottleneck.” With the number of concurrent flows improved by up to 50x, and the number of vNICs by up to 40x, the bottleneck s now the VM kernel stack, the researchers wrote.

Dai’s Forrester said that Nezha’s stateless offloading and cluster-wide pooling design is superior to solutions being pursued by rival cloud service providers.

Separately, Alibaba’s cloud computing division has also been working on another software update that will enable it to provide better network performance for AI workloads.

Called Alibaba Stellar, the update bypasses the scalability and stability challenges posed by its current remote direct access memory (RDMA) stack for high-performant workloads or AI services, providing a 15x improvement in container initialization time and improving LLM training speed by up to 14%, the researchers report.

Stellar, according to Dai, is more scalable and stable for large-scale AI clusters when compared to the support stack for RDMA used by AWS, Azure, or Google Cloud.

Other updates geared towards network improvement to be presented by Alibaba researchers at SIGCOMM this year includes Hermes, SkeletonHunter, and SkyNet.

SkeletonHunter is aimed at diagnosing and localizing network failures in containerized large model training, while Hermes is aimed at addressing the challenge of balancing connection distribution in multicore Layer-7 cloud load balancers.

Alibaba Cloud has deployed Hermes on 100,000 CPU cores, the researchers reported, with the result that it reduced daily worker hangs by 99.8% and lowered the unit cost of Layer-7 load-balancing infrastructure by 18.9%.

Anirban Ghoshal

Anirban is an award-winning journalist with a passion for enterprise software, cloud computing, databases, data analytics, AI infrastructure, and generative AI. He writes for CIO, InfoWorld, Computerworld, and Network World. He won the 2024 Silver Azbee Award for Best News Article in the Technology category. He has a post-graduate diploma in journalism from the Indian Institute of Journalism and New Media. Have a tip, scoop, or insight involving AI, cloud, databases, ERP, or enterprise software? Reach him securely on Signal at Ghoshal_CloudaiSaaSscoop.99

More from this author