H200
UK
GPUAAS.COM · WHOLESALE GPU NETWORK / A HOSTED·AI SERVICE ◆ CAPACITY AVAILABLE · 20+ PARTNERSQUOTES < 24HREV 2026.09
+
+
◆
H200 cost per million tokens
◆ AVAILABLE

H200
cost per token
, at
wholesale price.

Real H200 rental cost-per-million-token economics, and when the memory premium actually pays for itself, at
~30% less than hyperscale. Quotes in under 24 hours.

HGX GPU node
GPU generations
4
Architectures
Hopper + Blackwell
Vetted partners
20+
Quote turnaround
24 hrs
Commitment
Short / long
QUOTES IN UNDER 24 HOURS VETTED PARTNERS WORLDWIDE SHORT OR LONG TERM COMMITMENT DIRECT OPERATOR CONTRACTS CAPACITY AVAILABLE NOW PLACEMENT YOU SPECIFY
◆ THE SHORT ANSWER

H200's higher hourly rate doesn't automatically mean a higher cost per token; it depends entirely on whether the memory and bandwidth advantage translates into fewer GPUs or faster throughput for your specific workload. For a model that needed two H100s in tensor-parallel purely to fit KV cache, moving to a single H200 can lower total cost per token despite the higher per-GPU rate, since you're no longer paying for two GPUs. For a model that already fits comfortably on one H100, H200 typically raises cost per token for the same throughput, since compute is identical but the hourly rate is higher. H100 cost per token covers the same batching, quantization and engine levers that apply equally to H200.

+
01
PRICING

What actually drives H200 cost per token

H200 median on-demand rate runs $4.40/hr, a real premium over H100's $3.33/hr median. Whether that premium lowers or raises your actual cost per token depends entirely on whether it lets you avoid splitting a model across multiple GPUs. NVIDIA H200 pricing is quoted per enquiry; full NVIDIA H200 specs are available on request.

Market reference as of September 2026, quoted in USD. Cost-per-token figures shown are illustrative; actual figures depend heavily on model size, whether H200's memory avoids multi-GPU splitting, batch size, and quantization.
Wholesale rates through GPUaaS.com are quoted per enquiry and vary by commitment term, configuration and placement.

$0$2.50$5$7.50$10$12.50$15/GPU-HR
H100, single-stream BF16
One H100, per million tokens
$0.90
H100 x2 (tensor-parallel), FP8
Two H100s, model needs the split
$0.51
H200 x1, FP8 (same model)
One H200, no split needed
$0.30
H200, single-stream BF16
One H200, model already fit on H100
$1.19
◆ GPUaaS.com wholesale
Vetted partners · direct operator contract
quoted per enquiry
H200 COST PER MILLION TOKENS, BY SETUP
+
02
◆
What actually moves H200 cost per token

The levers that change H200 cost per token, and by how much.

The naive comparison, H200's $4.40/hr against H100's $3.33/hr, misses the actual question, which is total cost per token for the complete setup a given model needs. A model that fits comfortably on one H100 sees no benefit from H200 at all, since compute is identical between the two generations; the higher hourly rate directly raises cost per token for the exact same work. The calculation flips for models that need multi-GPU tensor-parallel splitting on H100 purely to fit KV cache or optimizer state, not because the compute demands it: two H100s at their combined hourly rate can cost more in total than one H200, and if the single H200 achieves comparable throughput, its cost per token comes out lower despite the higher per-GPU rate. The practical rule is straightforward: check whether your model needs multi-GPU splitting on H100 for memory reasons before assuming H200's premium is automatically a worse deal.

/01

Removing unnecessary GPU splits

For models needing tensor-parallel splitting across two H100s purely to fit memory, one H200 can lower total cost per token despite its higher rate.
avoids split · lower total cost · memory-bound
/02

No benefit when H100 already fits

For models that already fit on one H100, H200 raises cost per token for the same work, since compute is unchanged but the rate is higher.
no benefit · higher cost · compute unchanged
/03

Same optimization levers apply

Batching, FP8 quantization, and engine choice move H200 cost per token by the same multiples they do on H100.
batching · FP8 · same levers
/04

Compare total setup cost, not rate

Compare the full GPU-count cost, not just the hourly rate, since two H100s can cost more total than one H200 for the same model.
total cost · GPU count · real comparison
+
03
◆ LIVE NETWORK · 12 LOCATIONS

H200 capacity worldwide, in the location you need.

The hourly rate you pay for H200 capacity is one input into cost per token, and it varies meaningfully by region and provider. See H200 availability by country below.

Read the full guide to GPU cloud in this location →
4
GPU GENERATIONS
20+
VETTED PARTNERS
12
PLACEMENT OPTIONS
24h
QUOTE TURNAROUND
◆ USA◆ CAN◆ UK◆ DEU◆ FRA◆ NLD◆ UAE◆ SAU◆ IND◆ SGP◆ JPN◆ AUS
04
◆ COST COMPARISON

See how much you save at scale

Wholesale rates against cloud list price for a 64-GPU cluster.

CLUSTER SIZE
8 GPU Servers
64 × GPUS · 730 HRS/MO
ASSUMPTIONS · BLENDED $6.00/GPU-HR · INDICATIVE ONLY
SOURCEEST. MONTHLYVS GPUAAS
Retail cloud
On-demand list price · reserved discounts require lock-in
~$280k
+$84k
Direct datacentre negotiation
Long-term commitment · slow procurement cycle
~$230k
+$34k
◆ BEST VALUE
GPUaaS.com wholesale
Vetted partners · direct operator contract · quotes in 24 hours
~$196k
SAVE ~$84k/MO
Need single-GPU compute? packet.ai has you covered.
+
05
◆ HOW IT WORKS

A matchmaker, not a marketplace.

We connect you to our vetted partners. You contract directly with the operator running your nodes.

STEP 01/4
01

Tell us the requirement

GPU model, count, placement and timeline. Add workload detail if you have it.

STEP 02/4
02

We match capacity

We find vetted partners with capacity that fits, in the jurisdiction you need.

STEP 03/4
03

Quotes in 24 hours

Real quotes from partners who hold the capacity, not listings that may not exist.

STEP 04/4
04

Contract and provision

You contract directly with the operator. We smooth the provisioning process.

Get a quote
Request wholesale rates
in under 24 hours.

Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.

◆Quotes in under 24 hours
◆Direct contact with operators
◆Vetted partners, matched to your requirement
◆20+ vetted providers · 10 regions
1
ESSENTIALS
2
OPTIONAL
Contact
Full Name *
Business Email *
Organization *
Preferred Location *
Your Region *
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU Requirements
GPU Model *
PRE-SELECTED
Number of GPUs *
Individual GPU count. 1 node = 8 GPUs.
Get the Best Deal→
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
+
06
◆ FAQ

Frequently Asked Questions

Q1
Does H200 ever lower cost per token compared to H100?

It can, but only in specific cases: when H200's extra memory lets you avoid splitting a model across two or more H100s, the savings from needing fewer GPUs can outweigh H200's higher hourly rate. For models that already fit comfortably on one H100, H200 raises cost per token since compute is identical but the rate is higher.

Q2
How do I calculate whether H200 saves money for my specific model?

Compare total cost across the full setup, not just the hourly rate: a model needing two H100s at $3.33/hr each costs $6.66/hr in GPU rental, versus one H200 at $4.40/hr. If the H200 setup achieves comparable or better throughput on one card, its total cost per token can be lower despite the higher per-GPU rate.

Q3
Do the same batching and quantization levers apply to H200 cost per token?

Yes, identically. Batching multiple requests, using FP8 instead of BF16, and choosing an efficient serving engine all move cost per token by similar multiples on H200 as they do on H100, since these are software-level optimizations independent of which Hopper-class GPU you're running.

Q4
What's the clearest case where H200 wins on cost per token?

The clearest case is a model that needs multi-GPU tensor-parallel splitting purely to fit KV cache or optimizer state on H100, not because compute demands it. Removing that unnecessary GPU split by moving to a single H200 is where the cost-per-token math tends to favor H200.

Q5
Is H200 ever a bad choice for cost per token?

No. For models that comfortably fit on a single H100 already, H200 provides no throughput benefit since compute is identical between the two generations, so the higher hourly rate directly raises cost per token for the same work.

Q6
How do I decide between H100 and H200 for my specific cost-per-token situation?

Start by determining whether your model needs multi-GPU splitting on H100 purely for memory reasons. If it does, price the total multi-H100 setup against a single H200 running the same model. If your model already fits on one H100, H200 offers no cost-per-token advantage.