Blog ▸ Inference Is Now the Majority of GPU Spend. Procurement Has Not Caught Up.
GPU Infrastructure
Training spend has a calculable ceiling. Inference does not. Three variables move cost per token more than GPU choice does, and none appear on a rate card.
Inference Is Now the Majority of GPU Spend. Procurement Has Not Caught Up.
GPUaaS.com Team
GPU Infrastructure
September 17, 2026
No items found.
Inference took $23.3 billion of a $42 billion AI cloud market in 2026, passing training for the first time.
The number matters less than what it implies about cost shape.
Key takeaways
Training spend has a calculable ceiling. Inference does not, because it buys sustained capacity against load that never stops
Three variables move effective cost per token more than GPU choice does: utilization, context length, and serving configuration
None of those three appear on a rate card, which is why an inference quote priced on hardware answers the smallest question
Unit cost falls roughly 50% a year while volume rises faster, so cheaper tokens produce larger bills
Model tiering can cut cost per token 80 to 88%, a larger saving than any hardware negotiation produces
◆ TWO DIFFERENT COST SHAPES
One has a ceiling. The other does not.
Training spend has a ceiling you can calculate before you start. Pick a model size, a token budget, and a cluster, and the total is knowable. The run ends. The bill stops.
Inference has no such ceiling, because the quantity being bought is not compute hours but sustained capacity against load that does not stop.
That distinction explains why inference budgets break in ways training budgets do not, and the mechanism is visible in three numbers from the same fleet.
◆ WHAT ACTUALLY SETS COST PER TOKEN
Variable
Measured swing
On a rate card?
Utilization
$0.21 to $15.25 per M tokens
No
Context length
8x memory, 64x compute at 128K vs 16K
No
Batch configuration
26.6x throughput, batch 1 to 64
No
GPU model
Smaller than any of the above
Yes
◆ THE THREE NUMBERS
Utilization, context, configuration
The first is utilization. On identical H100 hardware, effective cost per million output tokens spans $0.21 to $15.25 depending on offered load, a penalty of up to 36 times near idle. A training cluster runs saturated by construction. An inference fleet runs at whatever concurrency the product happens to have that week, and the cost per token moves with it. The full picture is in how cost per token scales with concurrency.
The second is context length. A 128K window needs eight times the KV cache of 16K and, under full attention, sixty-four times the attention computation. A 70B model at 128K carries roughly 42.9 GB of KV cache for a single user. Nothing about that appears in a GPU count, and the arithmetic is worked through in what 128K windows do to your bill.
The third is configuration. Moving batch size from 1 to 64 on a single A100 changes throughput by 26.6 times. INT4 quantization cuts memory 75%, which raises the batch ceiling further because batch size is capped by memory rather than compute. Those levers are covered in three levers that beat new hardware.
55 cents
of every AI cloud dollar now goes to inference rather than training, the first year that has been true, and the procurement process was built for the other side of it
Gartner AI cloud market forecast, 2026
◆ WHY HARDWARE PRICING FITS ONE AND NOT THE OTHER
The smallest question in the stack
Those three variables move effective cost per token by more than the choice of GPU does. None of them appear on a rate card.
A training quote can be priced on hardware because the workload is fixed and known. An inference quote priced the same way is answering the smallest question in the stack.
◆ AGENTIC WORKLOADS WIDEN THE GAP
Cheaper tokens, larger bills
Agentic AI consumes 5 to 30 times more tokens per task than a standard chatbot interaction. One request becomes a plan, several tool calls, a critique, and a revision, each with its own prefill and decode. Reasoning models add long internal chains the user never sees and always pays for.
Unit cost falls roughly 50% a year. Volume rises faster. Token consumption is on track to multiply 24 times between 2026 and 2030, which is why cheaper tokens produce larger bills.
◆ THE SUPPLY SIDE MOVED FIRST
Memory is a third of hyperscaler spending
Memory now takes roughly 30% of hyperscaler AI data-center spending, four times its 2023 share, because inference is memory-bandwidth bound in ways training is not. Custom silicon aimed at inference is growing at a 44.6% compound annual rate. Hardware vendors are building for the workload that now dominates demand.
Inference ran at roughly 33% of compute in 2023, 50% in 2025, and 66% in 2026. Lenovo, counting shipped units rather than modelling from outside, puts the eventual split at 80/20 favouring inference. The direction is not in dispute even where the exact share is.
◆ WHAT CHANGES FOR ANYONE BUYING CAPACITY
Four adjustments
Separate the budget lines. Training is capital against a defined deliverable and inference is opex against a load curve, and a single line hides which one is growing.
Size against the load curve rather than the peak or the average. Reserved capacity is worst at low concurrency and best at high, with the crossover around 10 to 20 concurrent users in measured data, so the shape of demand decides the contract type.
Put facilities questions in the same document as model questions. Inference racks drawing 50 to 150 kW make power and cooling a procurement variable rather than a site detail.
Treat model tiering as a purchasing decision. A fine-tuned small model reaching 90% of frontier quality on a specific task can cut cost per token 80 to 88% against sending every query to a frontier model, which is a larger saving than any hardware negotiation will produce.
Across 1,192 organisations and $83 billion in cloud spend, AI workloads account for 18% of cloud spend at AI-forward enterprises, and 98% of practitioners actively manage AI spend.
Most of that management happens at the invoice. The variables that actually set the invoice sit one layer down, in utilization, context length, and serving configuration.
A quote that asks only for a GPU count is pricing the part that varies least.
Get quoted on the variables that set the bill.
Load curve, context length, serving configuration. No buyer fees. For single GPUs, packet.ai handles self-serve access with 24/7 human support.
Yes. Inference took $23.3 billion of a $42 billion AI cloud market in 2026, about 55 cents of every AI cloud dollar. Compute share ran roughly 33% in 2023, 50% in 2025, and 66% in 2026, and vendor order books point the same direction.
Training has a ceiling you can calculate from model size, token budget and cluster, and the run ends. Inference buys sustained capacity against load that does not stop, so the total depends on utilization, context length and serving configuration rather than on a fixed workload.
Volume rises faster than unit cost falls. Agentic workloads consume 5 to 30 times more tokens per task than a standard chatbot interaction, and total token consumption is on track to multiply 24 times between 2026 and 2030.
Load curve at peak and median, target context length, and serving configuration. Those three move effective cost per token more than the GPU model does, and a quote priced on hardware alone is pricing the part that varies least.
Model tiering. A fine-tuned small model reaching 90% of frontier quality on a specific task can cut cost per token 80 to 88% against routing every query to a frontier model, which exceeds what any hardware negotiation produces.
Last reviewed: 18 September 2026. AI cloud market split and agentic token multipliers from Gartner's 2026 forecast, reported via TechTimes, August 2026. Compute share series from Deloitte's 2026 TMT Predictions via AgentMarketCap. Shipped-unit split from Lenovo via Computerworld, January 2026. Token consumption growth from Goldman Sachs Research via Telnyx. Memory share and custom silicon growth rate from AgentMarketCap, April 2026. Cloud spend and practitioner figures from the FinOps Foundation's 2026 State of FinOps report. Model tiering savings from task-specific model benchmarks via TechTimes. Utilization, context and configuration figures from the GPUaaS analyses linked above. Browse current GPU cluster availability on GPUaaS.com.