With GPUs in short supply, cloud computing providers are increasingly turning to custom chips for specific workloads to deliver more cost-effective computing.
GPUs have become increasingly important for several large software firms such as AWS, Google, and OpenAI, as the demand for generative AI continues to grow steadily.
The service, currently in preview, will allow enterprises to run their real-time AI inferencing applications serving large language models on Nvidia L4 GPUs inside the managed service.
Google reportedly offered CISPE €14 million in cash and €455 million in software licenses to maintain its complaint with the EU, but ultimately CISPE decided to settle with Microsoft.
The expansion will be aided by technology from US-based electric vehicle manufacturer Tesla, according to reports from Chinese news outlets.
The data centers in Mumbai and Sydney are expected to close in July and September respectively, a notice from the company read.
The layoffs impacted employees working in Azure for Operators and Mission Engineering teams.
The strategy to lower prices may not only help Alibaba undercut competition from larger hyperscalers in emerging markets but also have a more positive effect on its image as a Chinese provider, experts say.
Azure Compute Fleet, a new service designed to simplify Azure provisioning, is also among the cloud infrastructure updates unveiled at Microsoft's Build 2024 event in Seattle.
Trillium, the sixth iteration of Google’s Tensor Processing Unit (TPU), is nearly five times more efficient than its predecessor, TPUv5, in peak compute performance and memory bandwidth, Google said.