For the first two years, the story was simple: GPU prices could only go down. Prices of H100 fell from roughly $8 per hour in 2023 to about $1.70 per hour on one-year contracts by late 2025. Everyone in the market saw the future – the oversupply is coming, and before long, computation will become commoditized.
But things changed dramatically in 2026.
In the middle of spring, Anjney Midha, founder of AMP, put it bluntly: “We are living through the COVID of compute, and all the toilet paper is gone.” There was no available capacity left on virtually all generations of GPUs. The trend in pricing has reversed, not in the direction of lower but rather in the direction of sky-high prices. And those who adopted early benefited from their insight.
As the crunch was being observed live among 150+ providers, let us tell you exactly how the crunch impacted your pricing.
The AI Compute Threshold Report
We analyzed pricing from 150+ GPU cloud providers to find the exact threshold where an AI startup's OpenAI API bill eclipses the cost of a dedicated H100 cluster.
Read the Full ReportThe crunch nobody wanted to accept
The structure only increased the impact. Based on the efficiency claims DeepSeek made in 2025, everybody became sure that there was an overinvestment in the field of compute and that it would give cheap compute forever.
But in early 2026, it seemed that it was not easy. SemiAnalysis described the hunt for GPU compute as being like trying to book airplane tickets on the last flight out: high prices, almost no availability. Those who rented their compute instances on the spot did not return them, no matter how the prices were growing.
The most surprising thing about all this is the increase in the prices for the old-generation compute. It is simply not supposed to happen this way with the arrival of the next generation.
What the 2026 GPU capacity crunch did to rental prices
The numbers tell it simply-
| GPU | Mid-2026 rate | The move |
|---|---|---|
| H100 (1-year contract) | ~$2.35/hr | up ~40% since October 2025 |
| H200 | ~$2.80/hr | climbing steadily |
| B200 (Blackwell) | ~$4.90 to $5.64/hr | up 48% in 60 days, roughly doubled |
| B300 | ~$5.29/hr | new premium tier |
| GB300 | ~$5.88/hr | scarcest and priciest |

Our story starts with the H100, which is just the beginning. The price for 1-year contracts on the H100 rose by 40%, reaching the level of $2.35 per hour in March compared to $1.70 per hour in October. An increase in prices for the outdated processor of two generations back. However, it happens only if there is more demand than supply for the entire stack, not just the top end.
The madness at the top end was even more impressive. The spot rental price on the Blackwell was as high as $4.08 per GPU hour in April, increasing 48% from $2.75 two months earlier, which attracted so much attention that it was mentioned in the Bloomberg Terminal. In July, the price of the B200 tier rose from $3.81 to $5.64. Meanwhile, the price on the Blackwell services increased nearly three times, compared to the previous record of the H100 processor.
Thus, in conclusion, we can say that prices have doubled since January.
What made it happen: three bottlenecks, a stampede, and everything else going wrong.
But that’s not one bottleneck. That’s three of them happening simultaneously, alongside with the increase in demand.
Packaging. Advanced GPUs need to have the memory attached through CoWoS technology provided by TSMC. And there is no room left whatsoever, none, zero, the calendar is completely booked up at least until mid-2027, which is the absolute maximum limit of accelerators to be physically manufactured. Nothing can be done about it.
Memory. Even more crucial than packaging. Samsung, SK Hynix, and Micron have started allocating capacities for the production of memory not for their standard DDR5 and NAND, but for HBM required for AI processors, and HBM3e prices have increased by around 20% in terms of 2026 contracts. AI initiatives are projected to consume up to 40% of global DRAM output. High prices of memory lead to high prices of the entire GPU, regardless of whether the manufacturer had some memory in stock.
Power and cooling. Blackwell GPUs require lots of power and liquid cooling infrastructure, which is not always provided by modern datacenters. Some hyperscalers have Blackwell GPUs already bought, but they are just unable to use them due to insufficient power and memory.
And then came the stampede. Demand surged for inference, not just training. NVIDIA quietly moved its wafer production from H100 to the more profitable Blackwell, which created an even scarcer supply of H100. On top of that, the huge hyperscaler orders that were worth billions of dollars exhausted the rest of NVIDIA’s allocation.
What really stings: the rent became more difficult, not just expensive
The war was won on price through pricing power. Short-term contracts were put secondary or even pre-emptable, payment was made upfront, and access to capacity required financial commitment. CoreWeave has hiked spot rates more than 20% since December 2025 and started pressuring small customers to enter three-year contracts.
Here is the way the crunch works. It’s not just about hourly prices. The inflexible commitment-oriented approach to cloud GPU rental that used to be the unique feature of the technology is slowly taken away from everyone without any forward commitment contract.
The impact went even to the labs. Outages occurred due to shortages at Anthropic, and OpenAI reduced its product portfolio in an attempt to decrease inference loads. And since they are squeezed, we all are squeezed.
It even got into your gaming GPU
As a quick aside, as it serves as evidence of just how deeply they hurt. As more of this memory is taken up by data-center equipment, less of it is available to consumer products. The RTX 5090 has seen roughly twice the cost in the aftermarket, from around $2,000 MSRP to over $4,000.
And if part shortages for data centers can raise the price of a gaming GPU, then…
But when will this end?
Short answer: not soon. But there is a real debate.
The bearish sentiment regarding the situation centers around timing. HBM and CoWoS shortages are going to last at least until the first six months of 2027, whereas H100 and H200 prices are not going to drop until additional memory capacity becomes available somewhere around the end of 2026 and beginning of 2027.
Unlike their rivals, Silicon Data claims that the spike in price of H100 was due to the structure issues and that prices are bound to drop once Blackwell becomes dominant in high end and H100 is going to dominate the middle segment.
It could be true for both. The frontier tier stays tight and expensive while previous-gen hardware slowly loosens up. In any case, counting on cheap computing power in 2026 is ill-advised.
Getting out of the crunch without overspending
That does not mean that the crunch implies overpaying for compute from the first offer made available to you. It means that there is now an increased disparity between providers.
- Consider both pricing and availability factors. What good will the lowest price listing do if there is nothing in stock? Under stress conditions, the second-lowest price provider that does not lack inventory is the clear winner.
- Find suppliers that place emphasis on on-demand offerings. The hyperscaler capacity is allocated to large enterprise customers first of all. On-demand cloud services specializing in GPUs provide exactly what they have to offer: on-demand capacity.
- Be creative about your capacity needs. FP8 or INT4 quantization techniques can reduce the needed capacity of the model in half, thus reducing the need for 8 H100s to 4 H100s. That saves you half the capacity and leaves you with the same results in an environment where each available GPU matters.
- Spot for tasks that can tolerate interruptions and reservations if possible. Spot pricing is more economical when you can afford interruption. Reservation gets you the pricing tier before the next influx of requests.
- Keep your options open with regard to hardware type and geographic location. Where one type of GPU is unavailable, another type or location could be the answer.
Do not allow the squeeze to dictate your choice for you
At such a close range of prices, from $1.38 per hour to $7.50 per hour, even the availability becomes as uncertain as ever. It’s something that you really do not want to risk.
With ComputeStacker, you can see current prices and availability from over 150 providers.
Frequently asked questions
The combination of three supply shortages , TSMC’s CoWoS capacity fully booked through mid-2027, the HBM memory shortage, and power/cooling limitations, alongside a surge in inference demand and hyperscaler forward bookings.
Roughly double since January 2026. H100 one-year contracts climbed about 40%, and Blackwell spot rates jumped 48% in 60 days, from $2.75 to $4.08. By July, the B200 tier was running $4.90 to $5.64.
Because the shortage was across the stack. NVIDIA prioritized wafers to more profitable Blackwell, the memory became scarce, and supply failed to catch up with demand for even two generations back hardware. The appreciation of older hardware is a clear indicator of the shortage.
Not earlier than late 2026 or early 2027, with the new HBM capacity coming into play. The frontier segment will most likely be the last one to experience oversupply, while previous-gen GPUs would ease off first.
Get personalised, no-commitment quotes from top AI infrastructure providers in under 2 minutes.



