The move could reshape cost planning for enterprises running large-scale AI workloads.
Amazon Web Services (AWS) has updated the pricing structure for some of its Elastic Compute Cloud (Amazon EC2) Capacity Blocks for ML offerings. Raised by approximately 15%, this move could affect enterprises planning large-scale machine learning workloads.
The Amazon Capacity Blocks allow customers to reserve access to accelerated compute resources for a future start date. The company allows customers to reserve accelerated compute instances in cluster sizes of one to 64 instances (512 GPUs or 1024 Trainium chips) for up to six months for running a broad range of ML workloads. However, the EC2 Capacity Blocks can be reserved only up to eight weeks in advance.
In June last year, AWS announced a price reduction of up to 45% for EC2 Nvidia GPU-accelerated instances across P4 and P5 instances. AWS did not immediately respond to a request for comment.
Prices climb across P5 Capacity Blocks
The Capacity Blocks include EC2 P6 instances that are accelerated by the latest Nvidia Blackwell GPUs, P5 instances complemented by Nvidia H100 and H200 Tensor Core GPUs, and P4 instances powered by Nvidia A100 Tensor Core GPUs.
Pricing for the p5e.48xlarge instance, featuring eight Nvidia H200 accelerators for the US East (Ohio) region, has increased from the effective hourly rate per instance (per accelerator) of $34.608 to $39.799.
For the p5en.48xlarge in the same region, the pricing has jumped from $36.184 to $41.612. This pricing remains the same across regions, including Stockholm, London, and Spain in Europe, and Jakarta, Mumbai, Tokyo, and Seoul in the Asia Pacific. However, customers in the US West (N. California) will now have to pay $49.749 instead of $43.26 for p5e.48xlarge and $52.015 instead of $45.23 for p5en.48xlarge.
The P6e pricing remains the same at $761.904 for 72 B200 accelerators in the Dallas Local Zone for p4d.24xlarge.
“EC2 Capacity Blocks for ML pricing are dynamic and vary based on supply and demand patterns, as described on the product detail page,” an AWS spokesperson said. “This price adjustment reflects the supply/demand patterns we expect this quarter. AWS’s commitment to not raise pricing on fixed pricing models like On Demand and Savings Plans remains unchanged.”
“The most defensible explanation is simply market-based pricing tied to supply and demand,” said Pareekh Jain, CEO at EIIRTrend & Pareekh Consulting. “As the demand for H100 and H200 GPUs outstrips supply, AWS is effectively applying a scarcity premium to guaranteed inventory. AWS is trying to recover higher infrastructure and capital costs from urgent capacity rather than overall capacity.”
Guaranteed GPU capacity becomes the new battleground
Guaranteed access to GPU clusters enables enterprises to de-risk AI infrastructure planning and build resilience against future supply volatility. Acknowledging the steep demand for high-end GPUs resulting in a shortage of Nvidia H100 and H200s, big clouds are increasingly offering guaranteed capacity to customers.
Other than AWS, Google and Microsoft also have similar offerings, but presented in more traditional reservation models and scheduling frameworks.
For instance, Google Cloud has introduced a calendar-based scheduling tool that lets customers reserve GPU capacity in fixed blocks ahead of time. “On paper, that looks a lot like what AWS is doing with Capacity Blocks. But the framing is different. Google is treating it as part of its broader resource scheduler, not a premium SKU. The guarantee is still there, but the pricing doesn’t feel as segmented or dynamic. It’s almost as if they’re using scheduling to compete, not price. And because they can also steer some workloads onto TPUs instead of GPUs, they’ve got a little more flexibility built into the system,” said Sanchit Vir Gogia, CEO and chief analyst at Greyhound Research.
Azure, on the other hand, leans more on regional capacity reservations. Gogia added that these allow customers to hold specific VM types in specific zones, but they tend to favour long-term planning and big enterprise commitments. They get guaranteed capacity, but often pay for the privilege of holding those resources, whether they use them or not. It’s a different kind of premium, less about price per hour, more about time and commitment.
Higher prices, limited immediate fallout
Experts say these offerings represent a smaller percentage of total cloud spend compared to general compute. However, they represent a disproportionately high share of strategic AI spend.
While AWS has been the first to come out to announce the price hike, Gogia believes the others might say it differently, but they’re working from the same playbook. Microsoft and Google also did not respond to the request for comment.
Jain explained that for most enterprises, the price hike is unlikely to trigger immediate migration. The EC2 Capacity Blocks typically account for a small portion of total GPU spend, and strong data gravity, entrenched MLOps stacks, compliance controls, and skills continue to anchor workloads on AWS. As moving mature ML training stacks off AWS remains complex and time-consuming, the impact will be felt more in new workloads.




