
Cost per token is the only GPU metric that maps to your actual bill, and almost nobody publishes it. Hourly rate tells you what a GPU costs; cost per token tells you what your product costs to run. The two frequently point in opposite directions, because a GPU at twice the hourly rate that delivers three times the throughput is a third cheaper per token. This page sets out how the arithmetic works, so you can run it against any provider's rates including ours.
The same hourly rate produces very different costs per token depending on context length, quantisation and whether the job fits on one GPU.
Placement affects your rate, because power cost and generation availability differ by market. Tell us what you need and you contract directly with the operator running the nodes.
We connect you to our vetted partners. You contract directly with the operator running your nodes.
GPU model, count, placement and timeline. Add workload detail if you have it.
We find vetted partners with capacity that fits, in the jurisdiction you need.
Real quotes from partners who hold the capacity, not listings that may not exist.
You contract directly with the operator. We smooth the provisioning process.
Tell us the essentials. We'll line up real quotes from our vetted wholesale partners, and you contract directly with the operator.