Vera Rubin
UK
{"@context":"https://schema.org","@graph":[{"@type":"Service","@id":"https://gpuaas.com/gpu/vera-rubin-llm-inference#service","name":"Vera Rubin for LLM Inference","provider":{"@type":"Organization","name":"GPUaaS.com","url":"https://gpuaas.com"},"serviceType":"GPU cloud infrastructure","description":"Vera Rubin NVL72 for LLM inference: specs, HBM4 bandwidth, early-access price and how it compares with GB300. Register interest, quoted per enquiry."},{"@type":"WebPage","@id":"https://gpuaas.com/gpu/vera-rubin-llm-inference#webpage","url":"https://gpuaas.com/gpu/vera-rubin-llm-inference","name":"Vera Rubin for LLM Inference","isPartOf":{"@type":"WebSite","name":"GPUaaS.com","url":"https://gpuaas.com"}},{"@type":"BreadcrumbList","@id":"https://gpuaas.com/gpu/vera-rubin-llm-inference#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https://gpuaas.com"},{"@type":"ListItem","position":2,"name":"GPU Cloud","item":"https://gpuaas.com/cluster"},{"@type":"ListItem","position":3,"name":"Vera Rubin for LLM Inference","item":"https://gpuaas.com/gpu/vera-rubin-llm-inference"}]},{"@type":"FAQPage","@id":"https://gpuaas.com/gpu/vera-rubin-llm-inference#faq","mainEntity":[{"@type":"Question","name":"Is Vera Rubin worth waiting for over GB300 for inference?","acceptedAnswer":{"@type":"Answer","text":"Only if your serving plan genuinely depends on the next generation's memory bandwidth and rack scale, and your timeline can absorb reservation-only access. GB300 is the same rack-scale design available now with a far deeper provider set, so for capacity this year GB300 or B300 are the more realistic choice."}},{"@type":"Question","name":"What are the Vera Rubin NVL72 specs for inference?","acceptedAnswer":{"@type":"Answer","text":"Vera Rubin NVL72 pairs 72 Rubin GPUs with 36 Vera CPUs (88 Arm Olympus cores each), 288GB of HBM4 per GPU at 22TB/s, connected via NVLink 6, with an estimated rack draw of 190-230kW. Some providers brand the Rubin GPU as \"H300\" for catalog continuity; it is the same silicon regardless of name."}},{"@type":"Question","name":"Why does HBM4 bandwidth matter for LLM inference?","acceptedAnswer":{"@type":"Answer","text":"Decode, the token-by-token phase that dominates serving time, is limited by memory bandwidth more than compute. HBM4 at 22TB/s per GPU, inside one NVLink 6 domain, targets exactly that for the largest models and longest contexts. Real gains depend on your model and serving engine, so validate on your own workload."}},{"@type":"Question","name":"What is Vera Rubin availability for inference today?","acceptedAnswer":{"@type":"Answer","text":"Broad on-demand access is not available yet. Vera Rubin is on early-access reservation terms through a small set of cloud partners, and confirmed regional availability varies by market. Register interest and we will confirm what is actually securable for your timeline."}},{"@type":"Question","name":"How much power does a Vera Rubin NVL72 rack need?","acceptedAnswer":{"@type":"Answer","text":"Roughly 190-230kW per rack is an estimate and sits well above GB300's roughly 135-140kW, so liquid cooling and secured power set real availability. Treat any in-country Vera Rubin supply as something to confirm directly, not assume."}},{"@type":"Question","name":"What is the Vera Rubin price per hour for inference?","acceptedAnswer":{"@type":"Answer","text":"Tracked early-access Vera Rubin pricing starts around $11.00/hr on a reservation-only basis as of September 2026, with one July 2026 industry analysis citing roughly $8.50/hr per chip in 3-year rental cost comparisons. The market is very thin, so Vera Rubin pricing through GPUaaS.com is quoted per enquiry and varies by commitment term and configuration."}}]}]}
GPUAAS.COM · WHOLESALE GPU NETWORK / A HOSTED·AI SERVICE ◆ CAPACITY AVAILABLE · 20+ PARTNERSQUOTES < 24HREV 2026.09
+
+
◆
Vera Rubin NVL72 for LLM inference
◆ AVAILABLE

Vera Rubin
for LLM inference
, at
wholesale price.

Register interest in Vera Rubin NVL72 early access, for inference on the largest models at rack scale, at
~30% less than hyperscale. Quotes in under 24 hours.

HGX GPU node
GPU generations
8
Architectures
Hopper + Blackwell + Vera Rubin
Vetted partners
20+
Quote turnaround
24 hrs
Commitment
Short / long
QUOTES IN UNDER 24 HOURS VETTED PARTNERS WORLDWIDE SHORT OR LONG TERM COMMITMENT DIRECT OPERATOR CONTRACTS CAPACITY AVAILABLE NOW PLACEMENT YOU SPECIFY
◆ THE SHORT ANSWER

Vera Rubin NVL72 (VR200) is NVIDIA's next rack-scale generation after Blackwell Ultra: 72 Rubin GPUs and 36 Vera CPUs in one NVLink 6 domain, with 288GB of HBM4 per GPU at 22TB/s of bandwidth. For inference, the memory bandwidth is the headline, since decode is bandwidth-bound and the rack keeps very large models and long-context KV caches inside one domain. It is early-access only today: tracked Vera Rubin price starts around $11.00/hr on reservation terms, and broad on-demand supply does not exist yet. Register interest and we will confirm what is actually securable for your timeline. GB300 for LLM inference, B300 and B200 are the available options for inference capacity this year.

+
01
PRICING

What Vera Rubin inference is expected to cost

Tracked early-access Vera Rubin pricing starts around $11.00/hr on a reservation-only basis, in a very thin market. Vera Rubin pricing is quoted per enquiry; full Vera Rubin NVL72 specs are available on request.

Market reference as of September 2026, quoted in USD. This is an early, thinly tracked market with few providers quoting publicly, so figures are directional. Real inference cost depends on model, precision and serving engine, so cost per token is the number to calculate.
Wholesale rates through GPUaaS.com are quoted per enquiry and vary by commitment term, configuration and placement.

$0$2.50$5$7.50$10$12.50$15/GPU-HR
Early-access rate, reservation-only
Tracked provider, reservation-only
$11.00
Industry reference (3-yr rental TCO)
Per-chip figure, July 2026 analysis
$8.50
Approx. premium band, high end
Directional, thin provider set
~$16.50
Hyperscaler on-demand (projected)
Likely floor once hyperscalers quote broadly
~$18.00
◆ GPUaaS.com wholesale
Vetted partners · direct operator contract
quoted per enquiry
◆ VERY THIN, EARLY MARKET: FEW PROVIDERS QUOTE YET
+
02
◆
Where Vera Rubin is expected to matter for inference

What Vera Rubin should serve well, and what to watch for.

Vera Rubin NVL72 is the next rack-scale step after Blackwell Ultra: 72 Rubin GPUs, 36 Vera CPUs, 288GB of HBM4 per GPU and one NVLink 6 domain. For inference, the case rests on memory bandwidth. Decode is bandwidth-bound, and HBM4 at 22TB/s per GPU is aimed squarely at it, while the rack keeps very large models and long-context caches off slower node-to-node links. None of that helps if you cannot get capacity: access today is early and reservation-only, supply is thin, and an estimated 190-230kW per rack means liquid cooling and secured power set where it can run. For inference you need this year, GB300 and B300 are the realistic options, and Vera Rubin is a deliberate plan for the largest serving workloads rather than a drop-in upgrade.

/01

Bandwidth for decode-bound serving

HBM4 at 22TB/s per GPU targets the bandwidth-bound decode phase, with 288GB per GPU keeping very large models and long-context caches in one domain.
HBM4 · 22TB/s · 288GB per GPU
/02

Rack-scale memory pool

72 Rubin GPUs share one NVLink 6 domain, so models and KV caches that span many GPUs avoid slower node-to-node links.
NVLink 6 · 72 GPUs · one domain
/03

Access today is early and narrow

Early access is reservation-only today, so Vera Rubin suits teams planning ahead, not those who need capacity now.
reservation-only · early access · plan ahead
/04

Rack-level power and supply

An estimated 190-230kW per rack means liquid cooling and secured power, so real in-country supply needs confirming.
190-230kW · liquid cooling · confirm supply
+
03
◆ LIVE NETWORK · 12 LOCATIONS

Vera Rubin capacity worldwide, in the location you need.

Vera Rubin availability is placement-sensitive and still forming: confirm real in-country supply before planning around it. See Vera Rubin availability by country below.

Read the full guide to GPU cloud in this location →
8
GPU GENERATIONS
20+
VETTED PARTNERS
12
PLACEMENT OPTIONS
24h
QUOTE TURNAROUND
◆ USA◆ CAN◆ UK◆ DEU◆ FRA◆ NLD◆ UAE◆ SAU◆ IND◆ SGP◆ JPN◆ AUS

Every location links to its own page. Click through for local pricing and specs.

04
◆ COST COMPARISON

See how much you save at scale

Wholesale rates against cloud list price for a 64-GPU cluster.

CLUSTER SIZE
8 GPU Servers
64 × GPUS · 730 HRS/MO
ASSUMPTIONS · BLENDED $6.00/GPU-HR · INDICATIVE ONLY
SOURCEEST. MONTHLYVS GPUAAS
Retail cloud
On-demand list price · reserved discounts require lock-in
~$280k
+$84k
Direct datacentre negotiation
Long-term commitment · slow procurement cycle
~$230k
+$34k
◆ BEST VALUE
GPUaaS.com wholesale
Vetted partners · direct operator contract · quotes in 24 hours
~$196k
SAVE ~$84k/MO
Need single-GPU compute? packet.ai has you covered.
+
05
◆ HOW IT WORKS

A matchmaker, not a marketplace.

We connect you to our vetted partners. You contract directly with the operator running your nodes.

STEP 01/4
01

Tell us the requirement

GPU model, count, placement and timeline. Add workload detail if you have it.

STEP 02/4
02

We match capacity

We find vetted partners with capacity that fits, in the jurisdiction you need.

STEP 03/4
03

Quotes in 24 hours

Real quotes from partners who hold the capacity, not listings that may not exist.

STEP 04/4
04

Contract and provision

You contract directly with the operator. We smooth the provisioning process.

Get a quote
Request wholesale rates
in under 24 hours.

Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.

◆Quotes in under 24 hours
◆Direct contact with operators
◆Vetted partners, matched to your requirement
◆20+ vetted providers · 12 locations
1
ESSENTIALS
2
OPTIONAL
Contact
Full Name *
Business Email *
Organization *
Preferred Location *
Your Region *
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
GPU Requirements
GPU Model *
PRE-SELECTED
Number of GPUs *
Individual GPU count. 1 node = 8 GPUs.
Get the Best Deal→
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
+
06
◆ FAQ

Frequently Asked Questions

Q1
Is Vera Rubin worth waiting for over GB300 for inference?

Only if your serving plan genuinely depends on the next generation's memory bandwidth and rack scale, and your timeline can absorb reservation-only access. GB300 is the same rack-scale design available now with a far deeper provider set, so for capacity this year GB300 or B300 are the more realistic choice.

Q2
What are the Vera Rubin NVL72 specs for inference?

Vera Rubin NVL72 pairs 72 Rubin GPUs with 36 Vera CPUs (88 Arm Olympus cores each), 288GB of HBM4 per GPU at 22TB/s, connected via NVLink 6, with an estimated rack draw of 190-230kW. Some providers brand the Rubin GPU as "H300" for catalog continuity; it is the same silicon regardless of name.

Q3
Why does HBM4 bandwidth matter for LLM inference?

Decode, the token-by-token phase that dominates serving time, is limited by memory bandwidth more than compute. HBM4 at 22TB/s per GPU, inside one NVLink 6 domain, targets exactly that for the largest models and longest contexts. Real gains depend on your model and serving engine, so validate on your own workload.

Q4
What is Vera Rubin availability for inference today?

Broad on-demand access is not available yet. Vera Rubin is on early-access reservation terms through a small set of cloud partners, and confirmed regional availability varies by market. Register interest and we will confirm what is actually securable for your timeline.

Q5
How much power does a Vera Rubin NVL72 rack need?

Roughly 190-230kW per rack is an estimate and sits well above GB300's roughly 135-140kW, so liquid cooling and secured power set real availability. Treat any in-country Vera Rubin supply as something to confirm directly, not assume.

Q6
What is the Vera Rubin price per hour for inference?

Tracked early-access Vera Rubin pricing starts around $11.00/hr on a reservation-only basis as of September 2026, with one July 2026 industry analysis citing roughly $8.50/hr per chip in 3-year rental cost comparisons. The market is very thin, so Vera Rubin pricing through GPUaaS.com is quoted per enquiry and varies by commitment term and configuration.