H100
UK
GPUAAS.COM · WHOLESALE GPU NETWORK / A HOSTED·AI SERVICE ◆ CAPACITY AVAILABLE · 20+ PARTNERSQUOTES < 24HREV 2026.09
+
+
◆
H100 80GB for LLM training
◆ AVAILABLE

H100
for LLM training
, at
wholesale price.

H100 80GB from vetted partners, sized to your model, framework and cluster topology, at
~30% less than hyperscale. Quotes in under 24 hours.

HGX GPU node
GPU generations
4
Architectures
Hopper + Blackwell
Vetted partners
20+
Quote turnaround
24 hrs
Commitment
Short / long
QUOTES IN UNDER 24 HOURS VETTED PARTNERS WORLDWIDE SHORT OR LONG TERM COMMITMENT DIRECT OPERATOR CONTRACTS CAPACITY AVAILABLE NOW PLACEMENT YOU SPECIFY
◆ THE SHORT ANSWER

Training and fine-tuning are memory-bound problems before they're compute-bound ones, and that's the main reason H100 remains the default GPU for LLM training in 2026. A model's optimizer state (Adam's momentum and variance terms) typically needs 4x the parameter count in memory on top of the weights themselves, so full fine-tuning of even a 7B model can need well over 100GB across GPUs before you've processed a single batch. H100's 80GB of HBM3 at 3,350 GB/s bandwidth (SXM5), combined with native FP8 training support, is why it's still the baseline target for PyTorch FSDP, DeepSpeed ZeRO and Megatron-LM. Full pretraining of large models needs many H100s networked over NVLink and InfiniBand; most teams fine-tuning an existing model use LoRA or QLoRA instead, which cuts trainable parameters by orders of magnitude and makes single-GPU or small-cluster fine-tuning realistic. H200 for LLM training reduces the sharding needed for models where 80GB genuinely isn't enough.

+
01
PRICING

What H100 training actually costs

H100 median on-demand rate runs $3.33/hr across 40+ tracked providers, ranging from $1.49 at the low end to $6.98 for specialist guaranteed-capacity providers. Training cost depends heavily on cluster size and run duration, so the hourly rate is a starting point, not the full picture.

Market reference as of September 2026, quoted in USD. Multi-GPU training cost also depends on cluster size, interconnect, and run duration, so the hourly rate is a starting point for estimating total training cost, not the full picture.
Wholesale rates through GPUaaS.com are quoted per enquiry and vary by commitment term, configuration and placement.

$0$2.50$5$7.50$10$12.50$15/GPU-HR
Market low, 40+ providers tracked
Cheapest tracked H100 SXM on-demand
$1.49
Median on-demand H100 SXM
Median across 40+ tracked providers
$3.33
Market high, specialist providers
Premium providers, guaranteed capacity
$6.98
Hyperscaler on-demand
What you pay without a broker
$12.29
◆ GPUaaS.com wholesale
Vetted partners · direct operator contract
quoted per enquiry
H100 MARKET RATES, AUGUST 2026
+
02
◆
Where H100 earns its keep in training

What H100 handles well for training, and what to watch for.

Training is dominated by a different bottleneck than inference: memory capacity during the backward pass, not just bandwidth during generation. Adam's optimizer state alone typically needs 4x a model's parameter count in memory, on top of the weights, gradients, and activations, which is why full fine-tuning of even mid-sized models can require sharding across multiple GPUs. H100's 80GB of HBM3 gives real headroom for LoRA and QLoRA fine-tuning on a single card, and its native FP8 support roughly doubles training throughput over BF16 once enabled, on top of being meaningfully faster than A100 per GPU at the same precision. For full pretraining or full fine-tuning of larger models, PyTorch FSDP or DeepSpeed ZeRO shard the model across a cluster of H100s connected via NVLink within a node and InfiniBand between nodes, trading added communication overhead for the ability to train models that don't fit on one card.

/01

Full pretraining and multi-GPU training

PyTorch FSDP or DeepSpeed ZeRO sharding a model's weights, gradients, and optimizer states across multiple H100s, networked over NVLink and InfiniBand.
FSDP · DeepSpeed · Megatron-LM
/02

Parameter-efficient fine-tuning

LoRA and QLoRA cut trainable parameters by orders of magnitude, making single-H100 fine-tuning of 7B-70B models realistic without a multi-GPU cluster.
LoRA · QLoRA · single-GPU
/03

FP8 mixed-precision training

Native FP8 training support roughly doubles throughput over BF16 once your training loop is set up to use it, cutting real GPU-hours per run.
FP8 · throughput · mixed precision
/04

Fitting large models in 80GB

Gradient checkpointing and activation offloading trade compute time for memory, letting larger batch sizes or longer sequences fit in 80GB.
checkpointing · memory · batch size
+
03
◆ LIVE NETWORK · 12 LOCATIONS

H100 capacity worldwide, in the location you need.

Training runs are often long-lived and data-residency sensitive, so where your cluster sits matters as much as the interconnect inside it. See H100 availability by country below.

Read the full guide to GPU cloud in this location →
4
GPU GENERATIONS
20+
VETTED PARTNERS
12
PLACEMENT OPTIONS
24h
QUOTE TURNAROUND
◆ USA◆ CAN◆ UK◆ DEU◆ FRA◆ NLD◆ UAE◆ SAU◆ IND◆ SGP◆ JPN◆ AUS
04
◆ COST COMPARISON

See how much you save at scale

Wholesale rates against cloud list price for a 64-GPU cluster.

CLUSTER SIZE
8 GPU Servers
64 × GPUS · 730 HRS/MO
ASSUMPTIONS · BLENDED $6.00/GPU-HR · INDICATIVE ONLY
SOURCEEST. MONTHLYVS GPUAAS
Retail cloud
On-demand list price · reserved discounts require lock-in
~$280k
+$84k
Direct datacentre negotiation
Long-term commitment · slow procurement cycle
~$230k
+$34k
◆ BEST VALUE
GPUaaS.com wholesale
Vetted partners · direct operator contract · quotes in 24 hours
~$196k
SAVE ~$84k/MO
Need single-GPU compute? packet.ai has you covered.
+
05
◆ HOW IT WORKS

A matchmaker, not a marketplace.

We connect you to our vetted partners. You contract directly with the operator running your nodes.

STEP 01/4
01

Tell us the requirement

GPU model, count, placement and timeline. Add workload detail if you have it.

STEP 02/4
02

We match capacity

We find vetted partners with capacity that fits, in the jurisdiction you need.

STEP 03/4
03

Quotes in 24 hours

Real quotes from partners who hold the capacity, not listings that may not exist.

STEP 04/4
04

Contract and provision

You contract directly with the operator. We smooth the provisioning process.

Get a quote
Request wholesale rates
in under 24 hours.

Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.

◆Quotes in under 24 hours
◆Direct contact with operators
◆Vetted partners, matched to your requirement
◆20+ vetted providers · 10 regions
1
ESSENTIALS
2
OPTIONAL
Contact
Full Name *
Business Email *
Organization *
Preferred Location *
Your Region *
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU Requirements
GPU Model *
PRE-SELECTED
Number of GPUs *
Individual GPU count. 1 node = 8 GPUs.
Get the Best Deal→
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
+
06
◆ FAQ

Frequently Asked Questions

Q1
What's the best GPU for LLM training in 2026?

For most training and fine-tuning work, H100 is still the practical default: the major frameworks (PyTorch FSDP, DeepSpeed, Megatron-LM) are built and tuned around it, and its native FP8 support materially speeds up training compared to A100. H200 and B200 are worth it specifically when you're memory-bound on very large models or context windows, since they add bandwidth and capacity rather than raw compute.

Q2
What's the difference between LoRA and QLoRA fine-tuning on H100?

LoRA (Low-Rank Adaptation) freezes the pretrained model weights and trains small additional rank-decomposition matrices, cutting trainable parameters by 100x or more. QLoRA adds 4-bit quantization of the frozen base model on top of that, cutting memory further. Both make it realistic to fine-tune a 7B-70B model on a single H100 instead of needing a multi-GPU cluster for full fine-tuning.

Q3
How much GPU memory does fine-tuning an LLM actually need?

As a rough rule, full fine-tuning of a dense model needs GPU memory equal to roughly 16-20x the parameter count when using Adam in mixed precision (weights, gradients, and optimizer states combined), before accounting for activations. A 7B model can therefore need over 100GB across GPUs for full fine-tuning, while LoRA fine-tuning of the same model can fit in under 24GB.

Q4
How does multi-GPU training work across H100 clusters?

Multi-GPU H100 training uses NVLink and NVSwitch for high-bandwidth communication within a node (900 GB/s per GPU) and InfiniBand for communication between nodes. Frameworks like PyTorch FSDP and DeepSpeed ZeRO shard the model, gradients, and optimizer states across GPUs so a model too large for one card's memory can still train, at the cost of added communication overhead.

Q5
How much faster is H100 than A100 for training?

H100's native FP8 support is the biggest factor: it roughly doubles achievable training throughput over BF16 once a model's training loop is set up to use it, on top of H100 already being meaningfully faster than A100 per GPU at the same precision. The combined effect is why H100 clusters can train comparable models in a fraction of the GPU-hours A100 clusters need.

Q6
What is gradient checkpointing and when should I use it for training?

Gradient checkpointing discards intermediate activations during the forward pass and recomputes them during the backward pass instead of storing all of them, trading extra compute time for significantly lower memory use. It's a standard technique for fitting larger batch sizes or longer sequences into a fixed amount of GPU memory during training.