Blog ▸ State of GPU Procurement Q4 2026: What Changed Since Q3
GPU Infrastructure
Median H100 on-demand hit $3.42 per GPU-hour in late September, up 15% year over year. The Blackwell glut did not arrive, because the bottleneck is memory.
State of GPU Procurement Q4 2026: What Changed Since Q3
GPUaaS.com Team
GPU Infrastructure
October 1, 2026
No items found.
In the week of 28 September 2026 the median on-demand H100 rented for $3.42 per GPU-hour across 40 providers. That is 15% higher than a year earlier.
The Blackwell glut that was supposed to collapse Hopper pricing has not arrived.
Key takeaways
Median H100 on-demand is up 15% year over year and the B200 benchmark is up 27.6% year to date
The bottleneck is memory rather than silicon. HBM and CoWoS packaging constrain supply, not GPU dies
The hyperscaler premium widened rather than closed: 58% on H100, 81% on B200, 141% on H200
Reservations run 29% below on-demand at one year and 49% at three, but three years now runs into the Rubin ramp
B200 capacity is still mostly sold through direct conversations six months after general availability
◆ THE SPREAD, LATE SEPTEMBER 2026
GPU
GPU clouds
Hyperscalers
Premium
H100
$5.00
$7.89
+58%
H200
$4.50
$10.85
+141%
B200
$7.88
$14.24
+81%
Median on-demand rates per GPU-hour. Source: GetDeploying price tracker via DataStorage State of GPU Cloud, September 2026
◆ WHAT CHANGED SINCE Q3
Prices went the wrong way
The expectation through Q2 and Q3 was that Blackwell volume would push Hopper into mid-tier pricing and drag the whole curve down. Hopper has indeed moved to mid-tier in role, but the price did not follow. The median H100 is up 15% year over year and the B200 benchmark is up 27.6% year to date.
Both things can be true at once. H100 rental has roughly halved since early 2024, and it has risen over the past twelve months. The decline happened earlier than the narrative suggests, and 2026 has been a partial reversal rather than a continuation.
The reason is memory. HBM supply and CoWoS packaging capacity at the foundry level constrain how many accelerators reach the market, and neither scales with die production. More Blackwell does not automatically mean more total compute when both generations draw on the same bottleneck.
+15%
year-over-year movement in the median on-demand H100 rate across 40 providers, in the quarter the market expected a Blackwell-driven glut
GetDeploying GPU price tracker, week of 28 September 2026
◆ THE SPREAD IS THE OPPORTUNITY
$1.49 to nearly $13 for the same silicon
An H100 rented anywhere from $1.49 to nearly $13 an hour in September depending on the provider. That is not a quality difference. It is margin, sales overhead and bundling.
The hyperscaler premium did not close during Q3. It sits at 58% on H100, 81% on B200 and 141% on H200, which is wide enough that the sourcing decision moves cost more than the hardware decision does.
H200 is the clearest anomaly. Specialist rates reached $3.99 an hour in October while hyperscaler list sits near $10.85, and specialist H200 now undercuts specialist H100 in several listings. Supply has caught up on that part faster than pricing has adjusted.
◆ LEAD TIMES FOR DIRECT PURCHASE
SKU
Current
2023 peak
H100 SXM
6-12 weeks
50+ weeks
H100 PCIe
4-8 weeks
30+ weeks
H200
8-14 weeks
Not shipping
B200
16-26 weeks
Not shipping
Cloud H100
Under 2 minutes
Waitlists
Secondary-market H100 from resellers currently runs 36-52 weeks, longer than new, driven by CoWoS and HBM constraints
◆ THE RESERVATION ARITHMETIC
$556,000 on a 64-GPU cluster, and the catch
September medians put a one-year reservation 29% below on-demand and a three-year term 49% below. On a 64-GPU H100 cluster at the $3.42 median, a year of on-demand costs about $1,917,389. The one-year reservation brings that to roughly $1,361,346, saving $556,043.
The three-year term gets the effective rate to $1.74 per GPU-hour, which looks decisive until the end date is checked. A three-year commitment signed now runs through 2029, across the Rubin ramp, on a forward curve that already points lower.
That is the central Q4 tension. Current pricing argues for locking in, and the generational calendar argues against it. Splitting the portfolio by term, with baseline load reserved and everything above it on demand, is the arrangement that does not require predicting which argument wins.
◆ B200 IS STILL OPAQUE
Six months after general availability
Most B200 capacity is still sold through direct sales conversations rather than listed on public rate cards. Published figures scatter accordingly, from $4.56 spot and $5.50 to $5.89 on-demand at specialists, up to around $14.24 at hyperscalers, with normalised multi-GPU pod pricing landing near $8.64.
For a buyer that means the listed number is a starting point rather than the price. Quoting B200 requires an actual conversation, and the spread between what is published and what is achievable is wider on this generation than on Hopper.
◆ WHAT TO DO WITH THIS IN Q4
Four moves the data supports
Check the sourcing channel before the hardware tier, because a 58 to 141% premium outweighs most tier decisions. Reserve the baseline and leave the variable portion on demand, since the forward curve makes a full three-year lock a directional bet rather than a saving.
Look hard at H200 while the anomaly lasts. Specialist rates at $3.99 against hyperscaler list near $10.85 is the widest relative gap in the market, and it is likely to narrow as pricing catches up with supply.
And treat secondary-market purchase with care. Reseller lead times of 36 to 52 weeks are longer than buying new, which makes used hardware a pricing reference rather than a practical source of capacity this quarter.
The headline for the quarter is that supply loosened and prices did not fall. Lead times normalised on Hopper, Blackwell reached volume, and the median rate still rose. A market where capacity is constrained by memory packaging rather than by die output does not behave the way a chip shortage ending usually behaves.
No. The median on-demand H100 reached $3.42 per GPU-hour in late September 2026, up 15% year over year, and the B200 benchmark is up 27.6% year to date. H100 has halved since early 2024, but that decline happened earlier and 2026 has been a partial reversal.
Because the constraint is memory rather than silicon. HBM supply and CoWoS packaging capacity limit how many accelerators reach the market, and both generations draw on the same bottleneck, so more Blackwell does not mean more total compute.
58% on H100, 81% on B200 and 141% on H200 against specialist GPU clouds on late-September medians. The full range on H100 ran from $1.49 to nearly $13 an hour depending on provider, which makes the sourcing channel a larger cost decision than the hardware tier.
It takes the effective rate to $1.74 per GPU-hour against a $3.42 median, but runs through 2029 across the Rubin ramp on a forward curve pointing lower. A one-year reservation at 29% below on-demand saves about $556,000 on a 64-GPU H100 cluster without the same exposure.
H100 SXM runs 6 to 12 weeks against 50-plus at the 2023 peak, H200 8 to 14, and B200 16 to 26. Cloud provisioning is effectively immediate. Secondary-market H100 from resellers runs 36 to 52 weeks, longer than buying new.
Last reviewed: 2 October 2026. Median rates, year-over-year movement, hyperscaler premiums and reservation arithmetic from GetDeploying's GPU price tracker across 40 providers, compiled in DataStorage's State of GPU Cloud September 2026 report. Provider-level rates from Spheron's GPU cloud pricing comparison, October 2026, and Thunder Compute's H200 price comparison, October 2026. The $1.49 to $13 range and B200 pricing opacity from Shattered's cloud GPU pricing analysis, September 2026. Lead times and supply constraints from Presenc's AI GPU supply and pricing research, and Spheron's H100 availability analysis. Memory and packaging bottleneck attribution from the same sources. Rates move frequently and should be checked before use in a model. Browse current GPU cluster availability on GPUaaS.com.