
Vera Rubin NVL72 (VR200) is NVIDIA's next rack-scale generation after Blackwell Ultra: 72 Rubin GPUs and 36 Vera CPUs in one NVLink 6 domain, with 288GB of HBM4 per GPU at 22TB/s of bandwidth. For inference, the memory bandwidth is the headline, since decode is bandwidth-bound and the rack keeps very large models and long-context KV caches inside one domain. It is early-access only today: tracked Vera Rubin price starts around $11.00/hr on reservation terms, and broad on-demand supply does not exist yet. Register interest and we will confirm what is actually securable for your timeline. GB300 for LLM inference, B300 and B200 are the available options for inference capacity this year.
Vera Rubin NVL72 is the next rack-scale step after Blackwell Ultra: 72 Rubin GPUs, 36 Vera CPUs, 288GB of HBM4 per GPU and one NVLink 6 domain. For inference, the case rests on memory bandwidth. Decode is bandwidth-bound, and HBM4 at 22TB/s per GPU is aimed squarely at it, while the rack keeps very large models and long-context caches off slower node-to-node links. None of that helps if you cannot get capacity: access today is early and reservation-only, supply is thin, and an estimated 190-230kW per rack means liquid cooling and secured power set where it can run. For inference you need this year, GB300 and B300 are the realistic options, and Vera Rubin is a deliberate plan for the largest serving workloads rather than a drop-in upgrade.
Vera Rubin availability is placement-sensitive and still forming: confirm real in-country supply before planning around it. See Vera Rubin availability by country below.
Read the full guide to GPU cloud in this location →Every location links to its own page. Click through for local pricing and specs.
Wholesale rates against cloud list price for a 64-GPU cluster.
We connect you to our vetted partners. You contract directly with the operator running your nodes.
GPU model, count, placement and timeline. Add workload detail if you have it.
We find vetted partners with capacity that fits, in the jurisdiction you need.
Real quotes from partners who hold the capacity, not listings that may not exist.
You contract directly with the operator. We smooth the provisioning process.
—
Need a different GPU generation? Each model available here has a dedicated page with full pricing and specifications.
Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.