H100
UK
GPUAAS.COM · WHOLESALE GPU NETWORK / A HOSTED·AI SERVICE ◆ CAPACITY AVAILABLE · 20+ PARTNERSQUOTES < 24HREV 2026.09
+
+
◆
H100 cost per million tokens
◆ AVAILABLE

H100
cost per token
, at
wholesale price.

Real H100 cost-per-million-token economics across engines, batch sizes and providers, at
~30% less than hyperscale. Quotes in under 24 hours.

HGX GPU node
GPU generations
4
Architectures
Hopper + Blackwell
Vetted partners
20+
Quote turnaround
24 hrs
Commitment
Short / long
QUOTES IN UNDER 24 HOURS VETTED PARTNERS WORLDWIDE SHORT OR LONG TERM COMMITMENT DIRECT OPERATOR CONTRACTS CAPACITY AVAILABLE NOW PLACEMENT YOU SPECIFY
◆ THE SHORT ANSWER

Cost per token is the number that actually determines whether an inference workload is economical, and it depends on far more than the GPU's hourly rental rate. The same H100 running the same model can produce wildly different cost-per-token figures depending on the serving engine (vLLM, SGLang, TensorRT-LLM all perform differently), the batch size (batching multiple requests together can cut cost per token by 3-4x versus single-stream serving), and the quantization used (FP8 roughly doubles throughput over BF16 on H100, directly halving cost per token for the same workload). A real example: H100 serving Llama 4 70B via vLLM runs about $0.45 per million tokens single-stream, dropping to roughly $0.12 per million tokens batched at 8 concurrent requests, a nearly 4x reduction from batching alone with no change in hardware or hourly rate. H200 cost per token can be lower in total when H200's memory avoids a multi-GPU H100 split.

+
01
PRICING

What actually drives H100 cost per token

The same H100 at the same hourly rate produces wildly different cost-per-token figures depending on setup. Moving from single-stream BF16 to batched FP8 cuts the real cost by roughly 7x, with no change to the hourly rate at all.

Market reference as of September 2026, quoted in USD. Cost-per-token figures shown are illustrative examples based on Llama 4 70B via vLLM; actual figures vary by model, engine, batch size, and quantization used.
Wholesale rates through GPUaaS.com are quoted per enquiry and vary by commitment term, configuration and placement.

$0$2.50$5$7.50$10$12.50$15/GPU-HR
Single-stream, BF16
vLLM, batch=1, per million tokens
$0.90
Single-stream, FP8
vLLM, batch=1, per million tokens
$0.45
Batched (8x), BF16
vLLM, batch=8, per million tokens
$0.24
Batched (8x), FP8
vLLM, batch=8, per million tokens
$0.12
◆ GPUaaS.com wholesale
Vetted partners · direct operator contract
quoted per enquiry
H100 COST PER MILLION TOKENS, BY SETUP
+
02
◆
What actually moves H100 cost per token

The levers that change H100 cost per token, and by how much.

Converting an hourly GPU rate into a cost-per-token figure requires knowing the actual throughput your specific setup achieves, which is where most naive cost estimates go wrong. Two teams running the identical H100 at the identical hourly rate can end up with cost-per-token figures that differ by 4x or more, purely based on serving engine choice and batch size. SGLang's RadixAttention, for instance, gives real throughput gains specifically when requests share common prefixes, which matters enormously for chat applications with repeated system prompts but less for one-off completions. FP8 quantization is close to a free throughput doubling on H100 since it was the first GPU generation to support it natively, meaning a workload still running BF16 is very likely leaving meaningful cost savings on the table without any hardware change at all.

/01

Batch size

Batching 8 concurrent requests versus single-stream serving can cut cost per token by roughly 3-4x on the identical hardware and hourly rate.
batching · 3-4x · concurrency
/02

FP8 vs BF16 quantization

H100 was the first GPU with native FP8 support, roughly doubling throughput over BF16 for near-free cost-per-token savings once enabled.
FP8 · quantization · 2x throughput
/03

Serving engine choice

vLLM, SGLang, and TensorRT-LLM perform differently by model and workload; SGLang's RadixAttention specifically helps shared-prefix chat traffic.
vLLM · SGLang · TensorRT-LLM
/04

Provider and hourly rate

The GPU's own hourly rate still matters, but it's a smaller lever than batching, quantization, and engine choice combined.
hourly rate · provider · wholesale
+
03
◆ LIVE NETWORK · 12 LOCATIONS

H100 capacity worldwide, in the location you need.

The hourly rate you pay for H100 capacity is one input into cost per token, and it varies meaningfully by region and provider. See H100 availability by country below.

Read the full guide to GPU cloud in this location →
4
GPU GENERATIONS
20+
VETTED PARTNERS
12
PLACEMENT OPTIONS
24h
QUOTE TURNAROUND
◆ USA◆ CAN◆ UK◆ DEU◆ FRA◆ NLD◆ UAE◆ SAU◆ IND◆ SGP◆ JPN◆ AUS
04
◆ COST COMPARISON

See how much you save at scale

Wholesale rates against cloud list price for a 64-GPU cluster.

CLUSTER SIZE
8 GPU Servers
64 × GPUS · 730 HRS/MO
ASSUMPTIONS · BLENDED $6.00/GPU-HR · INDICATIVE ONLY
SOURCEEST. MONTHLYVS GPUAAS
Retail cloud
On-demand list price · reserved discounts require lock-in
~$280k
+$84k
Direct datacentre negotiation
Long-term commitment · slow procurement cycle
~$230k
+$34k
◆ BEST VALUE
GPUaaS.com wholesale
Vetted partners · direct operator contract · quotes in 24 hours
~$196k
SAVE ~$84k/MO
Need single-GPU compute? packet.ai has you covered.
+
05
◆ HOW IT WORKS

A matchmaker, not a marketplace.

We connect you to our vetted partners. You contract directly with the operator running your nodes.

STEP 01/4
01

Tell us the requirement

GPU model, count, placement and timeline. Add workload detail if you have it.

STEP 02/4
02

We match capacity

We find vetted partners with capacity that fits, in the jurisdiction you need.

STEP 03/4
03

Quotes in 24 hours

Real quotes from partners who hold the capacity, not listings that may not exist.

STEP 04/4
04

Contract and provision

You contract directly with the operator. We smooth the provisioning process.

Get a quote
Request wholesale rates
in under 24 hours.

Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.

◆Quotes in under 24 hours
◆Direct contact with operators
◆Vetted partners, matched to your requirement
◆20+ vetted providers · 10 regions
1
ESSENTIALS
2
OPTIONAL
Contact
Full Name *
Business Email *
Organization *
Preferred Location *
Your Region *
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU Requirements
GPU Model *
PRE-SELECTED
Number of GPUs *
Individual GPU count. 1 node = 8 GPUs.
Get the Best Deal→
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
+
06
◆ FAQ

Frequently Asked Questions

Q1
How is H100 cost per token actually calculated?

Cost per token is the GPU's hourly rate divided by its actual achieved token throughput per hour, which depends on the serving engine, batch size, and quantization used. Two setups on the same GPU at the same hourly rate can have cost-per-token figures that differ by multiples based on these factors alone.

Q2
How much can batching reduce H100 cost per token?

In a real example, H100 serving Llama 4 70B via vLLM runs about $0.45 per million tokens single-stream, dropping to roughly $0.12 per million tokens batched at 8 concurrent requests, a nearly 4x reduction from batching alone with no hardware change.

Q3
Does FP8 actually lower cost per token on H100?

Yes, substantially. H100 was the first GPU generation with native FP8 support, and it roughly doubles achievable throughput over BF16 once enabled, which directly halves cost per token for the same workload assuming no other changes.

Q4
Which serving engine gives the lowest cost per token on H100?

It depends on the workload. TensorRT-LLM often wins on raw throughput if you can absorb a compile step, vLLM is a flexible default, and SGLang's RadixAttention specifically helps when requests share common prefixes, such as chat applications with repeated system prompts.

Q5
Is a lower hourly H100 rate always better for cost per token?

Not necessarily. A lower hourly rate on a poorly optimized setup (single-stream, BF16, no batching) can still produce a higher cost per token than a higher hourly rate running a well-tuned setup (batched, FP8, an efficient engine), since throughput matters as much as the rate.

Q6
How do I estimate cost per token for my own H100 workload?

Start by measuring your actual achieved tokens-per-second at your target batch size and quantization on your specific model, then divide the GPU's hourly rate by that throughput converted to tokens per hour. Published benchmarks are a starting point, but real throughput varies by model, hardware, and configuration.