Blog ▸ LLM Inference Cost on H100 in 2026: What It Actually Costs to Serve a Model at Scale
GPU Infrastructure
H100 cost per token depends on batch size and precision far more than the hourly rate. The real math, including why batching cuts cost 75% and where H100 still beats newer hardware.
LLM Inference Cost on H100 in 2026: What It Actually Costs to Serve a Model at Scale
GPUaaS.com Team
GPU Infrastructure
August 10, 2026
The world's most wanted GPU, NVIDIA B200 bare metal DC in US West - live on packet.ai →→(Access it from Bare metal CTA on top after login)
An H100 running Llama 4 70B single-stream generates roughly 95 tokens a second. Batched to 8 concurrent requests, 380. At a $2.85 hourly rate, that's $0.45 per million output tokens versus $0.12. Nearly 4x, no hardware change, no price change.
Most H100 cost conversations skip this number entirely. Annoying to explain to a finance team that wants one figure.
Key takeaways
Batching from 1 to 8 concurrent requests cuts H100 cost per token roughly 75% on Llama 4 70B, with zero hardware or price change
H100's throughput ceiling comes from memory bandwidth (3.35 TB/s), not raw compute, which is why single-stream inference looks expensive per token even at a competitive rate
Median H100 price sits around $8.97/hr across all listings, but the 25th-percentile price, arguably more honest, is $3.50. Spot pricing can run under $1/hr
Running inference on H100 costs roughly 11x more than on B300 for models that fit the newer hardware, per CloudZero's analysis, a real gap driving H100's secondary market decline
H100 still edges out newer tiers on throughput-sensitive 13B-70B parameter production APIs once FP8 quantization is in play
◆ THE BANDWIDTH CEILING
Memory, not compute, sets the limit
H100's throughput ceiling comes from memory bandwidth, not compute. 3.35 TB/s serves a 70B model at roughly 100 tokens a second single-stream. That ceiling is why H100 looks expensive per token even at a competitive hourly rate. The GPU isn't the bottleneck. Moving data through memory is.
◆ BATCH SIZE IS THE BIGGEST LEVER
75% cost reduction from software, not hardware
Batch size is the biggest lever on H100 cost per token. Single-stream leaves most of the GPU's compute idle while it waits on memory. Batching lets that same bandwidth serve more tokens per second without adding hardware. Batch=1 to batch=8 on Llama 4 70B cuts cost per token roughly 75%. Software configuration, not a faster chip.
FP8 does something similar from a different angle. On H100, FP8 roughly doubles throughput without changing GPU count. A team running FP16 and comparing cost per token against a competitor's FP8 numbers is comparing two operating points on the same hardware, not two different GPUs.
◆ WHAT "H100 RATE" ACTUALLY MEANS
Tier
Typical rate
Spot, specialty provider
Under $1.00/hr
25th percentile, all listings
~$3.50/hr
Median, all listings
~$8.97/hr
Hyperscaler on-demand
Up to 99% above specialty rate
Source: GPU Tracker 2026 Cloud GPU Statistics, 5,213 live listings across 54 providers
11x
how much more it costs to run inference on an H100 than on a B300 for models that fit the newer hardware, a real gap driving H100's secondary market decline
CloudZero H100 GPU Cost 2026 analysis
◆ THE GAP BEHIND THE SECONDARY MARKET DROP
Not a glitch, a real economic shift
Here's the part that matters for anyone deciding whether H100 makes sense at all right now. It costs roughly 11 times more to run inference on an H100 than on a B300, per CloudZero's analysis. Genuinely large gap. Not marketing exaggeration. Direct consequence of Blackwell's native FP4 support and higher memory bandwidth doing structurally more work per dollar on large models. Anyone still running production inference on H100 for models that fit comfortably on newer hardware needs to charge more per token than a competitor who migrated, or accept a thinner margin.
That gap is why H100 secondary market prices moved this fast. Cards that sold for $40,000 in late 2023 traded at $21,000 to $34,000 by Q2 2026. Forecasts point to another 10 to 20% decline as more H100 inventory returns to the secondary market from enterprises finishing multi-year contracts and migrating to B200. Not a market glitch. Blackwell making Hopper-class hardware economically obsolete for a specific slice of workloads, mostly inference on large models, while H100 stays genuinely competitive everywhere else.
◆ WHERE H100 STILL WINS, AND WHERE IT STOPS
The 13B-70B lane, until memory becomes the constraint
H100 still wins a specific lane even with B200 and H200 fully available. For throughput-sensitive production APIs serving 13B to 70B parameter models, H100 edges out A100 and, in a meaningful share of FP8 configurations, holds its own against newer tiers too. Not a nostalgia argument. A real result tied to how well H100's Transformer Engine handles quantized inference at that model size range. It's the reason more than one market analysis still calls H100 "broadly rentable and often the value choice," not a chip anyone should be embarrassed to run.
H100 stops winning exactly where memory bandwidth becomes the binding constraint rather than compute. A model needing more than 80GB to serve in FP16 forces quantization tradeoffs or a second GPU, either of which changes the cost math H100 was supposed to deliver. H200's 141GB fits many of those same models on one card, at a rate that has to be checked against what the second H100 would have cost, since the honest comparison is never one H100 against one H200.
Get a quote based on your actual throughput.
Not a generic hourly number. No buyer fees. For single GPUs, packet.ai handles self-serve access with 24/7 human support.
It depends heavily on batch size. Single-stream inference on Llama 4 70B costs roughly $0.45 per million output tokens at a $2.85/hr rate. Batched to 8 concurrent requests, that drops to roughly $0.12, nearly a 4x reduction with no hardware or price change.
H100's throughput is limited by memory bandwidth, not compute. Single-stream inference leaves most of the GPU's compute idle while it waits on memory. Batching multiple requests lets the same bandwidth serve more tokens per second, cutting cost per token by roughly 75% between batch=1 and batch=8 on a 70B model.
For throughput-sensitive 13B to 70B parameter production APIs, especially with FP8 quantization, H100 remains genuinely competitive. For models needing more than 80GB in FP16, or for maximum throughput on very large models, newer hardware pulls ahead, sometimes by a large margin, roughly 11x on inference cost against B300 for workloads suited to the newer chip.
The median H100 rate across all listings is around $8.97/hr, but the 25th-percentile rate, arguably a more honest read of what a real buyer pays, is $3.50/hr, and spot rates on specialty providers can run under $1/hr. Hyperscaler on-demand rates can run up to 99% higher than specialty provider rates for the identical chip.
If a model needs more than 80GB to serve in FP16, a single H100 can't fit it without quantization tradeoffs or a second GPU. H200's 141GB fits many of those same models on one card. Submit a spec through GPUaaS.com and get matched to the tier that actually fits, based on real throughput rather than a generic hourly rate.
Last reviewed: 11 August 2026. Throughput and cost-per-token figures from AITOT Inference Benchmark, July 2026, vLLM v0.25 at FP16. Pricing distribution data from GPU Tracker's 2026 Cloud GPU Statistics, 5,213 live listings across 54 providers. Inference cost comparison and secondary market data from CloudZero's H100 GPU Cost 2026 analysis and TBR Trade's Q2 2026 GPU Market Update. Get an H100 cluster quote on GPUaaS.com.