H100
UK
GPUAAS.COM · WHOLESALE GPU NETWORK / A HOSTED·AI SERVICE ◆ CAPACITY AVAILABLE · 20+ PARTNERSQUOTES < 24HREV 2026.09
+
+
◆ H100 80GB for LLM inference
◆ AVAILABLE

H100
for LLM inference
, at
wholesale price.

H100 80GB from vetted partners, sized to your throughput and context length, at
~30% less than hyperscale. Quotes in under 24 hours.

HGX GPU node
GPU generations
4
Architectures
Hopper + Blackwell
Vetted partners
20+
Quote turnaround
24 hrs
Commitment
Short / long
QUOTES IN UNDER 24 HOURS VETTED PARTNERS WORLDWIDE SHORT OR LONG TERM COMMITMENT DIRECT OPERATOR CONTRACTS CAPACITY AVAILABLE NOW PLACEMENT YOU SPECIFY
◆ THE SHORT ANSWER

Inference is the workload most GPU buyers underestimate, and the GPU for inference you pick has to earn its keep continuously rather than in bursts. The H100 remains the default choice for LLM inference in 2026, not because it's the newest card but because the software stack is mature around it. vLLM, SGLang and TensorRT-LLM all treat H100 as the baseline target, and hardware-native FP8 support (H100 was the first GPU to ship it) gives roughly double the throughput of BF16 for free once you enable it. On Llama 3.1 8B, SGLang reaches roughly 16,200 tokens per second on a single H100, about 29% ahead of vLLM's ~12,500. For a 70B model at FP8, the engine choice matters more than the hardware: TensorRT-LLM wins on raw throughput if you can absorb a compile step, vLLM is the safer default, SGLang's RadixAttention wins when requests share prefixes.

+
01
◆ PRICING

What H100 inference actually costs

H100 median on-demand rate runs $3.33/hr across 40+ tracked providers, ranging from $1.49 at the low end to $6.98 for specialist guaranteed-capacity providers. Real inference cost also depends on tokens per second at your model and batch size, which is why the engine matters as much as the rate.

Market reference as of September 2026, quoted in USD. Real throughput depends on model, quantization, context length and serving engine, so cost per token is the number to calculate, not the hourly rate alone.
Wholesale rates through GPUaaS.com are quoted per enquiry and vary by commitment term, configuration and placement.

$0$2.50$5$7.50$10$12.50$15/GPU-HR
Market low, 40+ providers tracked
Cheapest tracked H100 SXM on-demand
$1.49
Median on-demand H100 SXM
Median across 40+ tracked providers
$3.33
Market high, specialist providers
Premium providers, guaranteed capacity
$6.98
Hyperscaler on-demand
What you pay without a broker
$12.29
◆ GPUaaS.com wholesale
Vetted partners · direct operator contract
quoted per enquiry
◆ H100 MARKET RATES, AUGUST 2026
+
02
Where H100 inference earns its keep

What H100 serves well, and what to watch for.

Inference splits into two phases with different bottlenecks: prefill, which processes the prompt and is compute-bound, and decode, which generates tokens one at a time and is almost entirely memory-bandwidth bound. H100's 3,350 GB/s of bandwidth (SXM5) is why it still holds up against newer cards on the decode-bound half of the job that dominates real serving traffic. Against the previous generation, H100 runs roughly 1.7x faster than A100 on Llama-class models. Against H200, the newer card cuts time-to-first-token by around 40% thanks to higher bandwidth, which matters more for latency-sensitive chat than for batch throughput. 80GB of HBM3 fits most 70B-class models at FP8 with meaningful KV cache headroom, and 7 isolated MIG instances let you split one card across multiple smaller serving jobs without buying separate hardware.

/01

Production model serving

SGLang or vLLM serving Llama-class 8B-70B models at FP8, with MIG partitioning for multi-tenant traffic on a single card.
FP8 · vLLM · SGLang
/02

Long-context and RAG

80GB fits 70B at FP8 with real KV cache headroom for RAG and long-context serving, though H200's extra memory helps past very long windows.
80GB HBM3 · KV cache · RAG
/03

Multi-tenant via MIG

Partition one H100 into up to 7 isolated instances for smaller serving jobs, dynamically reconfigurable without a GPU reset.
MIG · 7 instances · isolated
/04

Migrating from A100

About 1.7x faster than A100 on comparable Llama-class inference, with native FP8 support A100 never had.
vs A100 · FP8 · upgrade path
+
03
◆ LIVE NETWORK · 12 LOCATIONS

H100 capacity worldwide, in the location you need.

H100 inference is placement-sensitive: prompts and responses are often personal data, so jurisdiction matters as much as latency. See H100 availability by country below.

Read the full guide to GPU cloud in this location →
4
GPU GENERATIONS
20+
VETTED PARTNERS
12
PLACEMENT OPTIONS
24h
QUOTE TURNAROUND
USA CAN UK DEU FRA NLD UAE SAU IND SGP JPN AUS
04
◆ COST COMPARISON

See how much you save at scale

Wholesale rates against cloud list price for a 64-GPU cluster.

CLUSTER SIZE
8 GPU Servers
64 × GPUS · 730 HRS/MO
ASSUMPTIONS · BLENDED $6.00/GPU-HR · INDICATIVE ONLY
SOURCEEST. MONTHLYVS GPUAAS
Retail cloud
On-demand list price · reserved discounts require lock-in
~$280k
+$84k
Direct datacentre negotiation
Long-term commitment · slow procurement cycle
~$230k
+$34k
◆ BEST VALUE
GPUaaS.com wholesale
Vetted partners · direct operator contract · quotes in 24 hours
~$196k
SAVE ~$84k/MO
Need single-GPU compute? packet.ai has you covered.
+
05
◆ HOW IT WORKS

A matchmaker, not a marketplace.

We connect you to our vetted partners. You contract directly with the operator running your nodes.

STEP 01/4
01

Tell us the requirement

GPU model, count, placement and timeline. Add workload detail if you have it.

STEP 02/4
02

We match capacity

We find vetted partners with capacity that fits, in the jurisdiction you need.

STEP 03/4
03

Quotes in 24 hours

Real quotes from partners who hold the capacity, not listings that may not exist.

STEP 04/4
04

Contract and provision

You contract directly with the operator. We smooth the provisioning process.

Get a quote
Request wholesale rates
in under 24 hours.

Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.

Quotes in under 24 hours
Direct contact with operators
Vetted partners, matched to your requirement
20+ vetted providers · 10 regions
1
ESSENTIALS
2
OPTIONAL
Contact
Full Name *
Business Email *
Organization *
Preferred Location *
Your Region *
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU Requirements
GPU Model *
PRE-SELECTED
Number of GPUs *
Individual GPU count. 1 node = 8 GPUs.
Get the Best Deal
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
+
06
◆ FAQ

Frequently Asked Questions

Q1
What's the best GPU for LLM inference in 2026?

For most production serving, H100 remains the default: the software stack (vLLM, SGLang, TensorRT-LLM) is built around it, hardware FP8 support roughly doubles throughput over BF16, and 80GB fits most 70B-class models with KV cache headroom. H200 and B200 win on very long context or where memory is the binding constraint.

Q2
How fast is H100 for LLM inference?

On Llama 3.1 8B, SGLang reaches roughly 16,200 tokens per second on a single H100, about 29% ahead of vLLM's ~12,500 (PremAI benchmark). On Llama 3.3 70B at FP8, throughput depends heavily on which of vLLM, SGLang or TensorRT-LLM you use.

Q3
Which inference engine is fastest on H100?

It depends on your constraint. TensorRT-LLM gives the best raw throughput if you can absorb a roughly 28-minute compilation step per model. vLLM is the safe, flexible default. SGLang's RadixAttention gives real gains specifically when requests share prefixes. TGI is now maintenance-only as of December 2025 and not a serious contender for new deployments.

Q4
How does batch size affect H100 inference throughput?

Batching multiple requests together on H100 raises total throughput substantially, since the GPU processes them in parallel rather than one at a time. The tradeoff is latency: a larger batch means each individual request waits longer for its turn, so the right batch size depends on whether you're optimizing for single-request speed or total requests served per second.

Q5
H100 vs H200 vs A100 for inference: what's the real difference?

H100 is roughly 1.7x faster than A100 on comparable inference and adds native FP8 support A100 lacks. H200 keeps the same compute but adds memory bandwidth, cutting time-to-first-token by about 40% versus H100, which matters most for latency-sensitive chat rather than batch throughput.

Q6
Can one H100 serve multiple models or customers?

Yes, via Multi-Instance GPU (MIG): H100 supports up to 7 isolated instances, each with dedicated memory and compute, and profiles can be changed dynamically without a GPU reset, which suits multi-tenant serving without needing separate hardware per workload.