
Cost per token is the hourly rate divided by sustained throughput, and RTX 5090's low rate makes it the cheapest starting point on the platform: a tracked median near $0.70/hr against $3.33/hr for H100. Whether that wins on cost per token depends on whether your model fits in 32GB. Models up to roughly 41B parameters at 4-bit run on one card, where the hourly saving carries straight through to cost per token, while models that need more memory force multiple cards or a larger GPU and the advantage narrows. Full RTX 5090 specs are available on request. H100 cost per token and H200 cover the same batching, quantization and engine levers that apply to RTX 5090.
RTX 5090's cost per token is driven by four levers more than by the headline rate: whether your model fits in 32GB, how aggressively you quantize, how well you batch and utilise the card, and the commercial term. Against H100, whose tracked median sits near $3.33/hr, RTX 5090 at a median near $0.70/hr starts far ahead, and for models up to roughly 41B parameters at 4-bit that hourly saving carries straight through to cost per token. The advantage narrows once a model needs more than one card or a context outgrows 32GB, which is where H100 and H200 earn their rate. RTX 5090 is a consumer-grade card, so operator terms and software licensing for hosted use vary and are worth confirming with the operator.
RTX 5090 is among the most widely distributed cards on the platform, so most countries have options. Confirm the operator and its terms for your location. See RTX 5090 availability by country below.
Read the full guide to GPU cloud in this location →Every location links to its own page. Click through for local pricing and specs.
Wholesale rates against cloud list price for a 64-GPU cluster.
We connect you to our vetted partners. You contract directly with the operator running your nodes.
GPU model, count, placement and timeline. Add workload detail if you have it.
We find vetted partners with capacity that fits, in the jurisdiction you need.
Real quotes from partners who hold the capacity, not listings that may not exist.
You contract directly with the operator. We smooth the provisioning process.
—
Need a different GPU generation? Each model available here has a dedicated page with full pricing and specifications.
Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.