Blog ▸ L40S for Inference and Fine-Tuning: The Tier Nobody Writes About
GPU Infrastructure
The L40S never tops a benchmark chart and has the lowest cost per token of the mainstream tiers. Why that happens, where it wins, and where the missing interconnect ends the conversation.
L40S for Inference and Fine-Tuning: The Tier Nobody Writes About
GPUaaS.com Team
GPU Infrastructure
August 26, 2026
No items found.
The L40S has the lowest cost per token of the mainstream inference tiers. $0.023 per million against the H100's $0.026.
It is also slower than an H100 at everything. Roughly 2x the latency on single-batch inference, 1.8x slower than an A100 on the same measure, and 864 GB/s of memory bandwidth against the H100's 3,350.
Key takeaways
Lowest cost per token of the mainstream tiers, because the 5.33x hourly price gap against H100 exceeds H100's 2.7-5x throughput advantage
L40S supports FP8 through the Transformer Engine. The A100 does not. That is why they are not really the same tier despite similar hourly rates
Same AD102 Ada Lovelace die as an RTX 4090, tuned for 24/7 data center operation, with 48GB GDDR6 ECC instead of 24GB
No NVLink, PCIe Gen4 only. PCIe 4.0 gives ~64 GB/s bidirectional, roughly 14x less than NVLink, so tensor-parallel work is out
As Blackwell ramps and A100/H100 secondary pricing softens, L40S is holding most of its value
Both of those are true at once, and that is the whole reason this card gets skipped in benchmark coverage. It never tops a chart. It wins the metric that decides a serving bill.
The mechanism is simple arithmetic. L40S runs roughly 5.33x cheaper per hour than an H100 SXM. The H100's throughput advantage on LLM serving runs 2.7x to 5x. When the price gap is larger than the throughput gap, the cheaper card produces more tokens per dollar even while producing fewer tokens per second.
The card itself is an odd position in the lineup. Same AD102 Ada Lovelace die as an RTX 4090, tuned for 24/7 data center operation, with 48GB of GDDR6 ECC where the 4090 carries 24GB. That memory doubling is the entire reason it is relevant for LLM inference rather than being a workstation part.
◆ THE FP8 LINE SEPARATES THE VALUE TIER
L40S has it. A100 does not.
Precision support is where it separates from the older value tier, and this is the fact most comparisons bury. The L40S supports FP8 through the Transformer Engine. The A100 does not. 724 TFLOPS at FP8, 362 at FP16, 91 at FP32. With FP8 quantization, 48GB behaves closer to 96GB for memory-bound serving, which changes what fits on one card substantially.
That single capability is why the L40S and A100 are not really the same tier despite similar hourly rates. An A100 tops out at BF16 for training and INT8 for inference. An L40S has the low-precision path that everything built after 2023 assumes.
Where the A100 still wins is bandwidth and capacity. 2 TB/s against 864 GB/s, and 80GB against 48GB. For memory-bandwidth-bound workloads and models in the 30B to 70B range at FP8, the A100 remains the better fit. The A100 also has MIG for multi-tenant partitioning, which the L40S lacks. And FP64 for double-precision scientific work is not a contest, the Ada die is weak there.
$0.023/M
L40S cost per million tokens against the H100's $0.026, despite the H100 being faster at every individual measurement
Cyfuture L40S vs A100 vs H100 benchmark comparison, 2026
◆ WHERE L40S IS THE RIGHT ANSWER
Workload
Why it fits
Image generation
SDXL at any standard resolution, Flux.1 to 1080p, comparable rate to H100 at 5.33x lower hourly cost
Embedding endpoints
BGE, GTE, sentence-transformers are light. Millions of embeddings per hour at low utilization
Fine-tuning to 30B
Near-Ampere speed at ~60% the hourly rate on RAG adapters, seq 128-256, batch 16-32
LLM serving 7B-40B
Batch sizes 1-8. Llama 3.1 8B around 336 tok/s
Not suitable
Tensor-parallel inference, distributed training, 70B+ on a single card
◆ WHAT IT COSTS
$0.55 to $1.66 depending on channel
Pricing across the market runs roughly $0.55 to $1.66 per GPU-hour depending on provider and channel, listed by around 33 providers. On packet.ai, dedicated single-tenant L40S runs $0.92 per hour with a monthly flat-rate cap at $604 regardless of hours used, which changes the calculation for anything running continuously.
◆ THE STRONGEST CASE IS THE LEAST DISCUSSED
Image generation and embeddings
Image generation is the strongest case and the least discussed. 48GB handles SDXL at any standard resolution and Flux.1 up to 1080p. L40S generates images at a rate comparable to an H100 while costing 5.33x less per hour, because most image generation deployments are latency-focused rather than batch-throughput-focused, and the H100's compute advantage only compounds at high parallel concurrency. SDXL images land around $0.00064 each on packet.ai rates, roughly 2.7x cheaper than the same output on H100.
Embedding endpoints are almost absurdly well matched. BGE, GTE, and sentence-transformers models are small relative to 48GB and computationally light. A single L40S serves millions of embeddings per hour at very low GPU utilization. Running those same endpoints on an H100 works fine and wastes most of the card.
Fine-tuning up to 30B parameters is achievable with appropriate optimization. For daily fine-tuning and RAG adapter work at sequence lengths of 128 to 256 and batch sizes of 16 to 32, the L40S delivers near-Ampere speed at around 60% of the hourly rate. Many small jobs rather than one large one is the pattern that fits. LLM serving in the 7B to 40B range at batch sizes of 1 to 8 is the general sweet spot, with Llama 3.1 8B running around 336 tokens per second.
◆ WHERE IT STOPS WORKING
No cost advantage fixes a missing interconnect
Where it stops working is worth being equally direct about. No NVLink, PCIe Gen4 only. PCIe 4.0 delivers roughly 64 GB/s bidirectional on a sixteen-lane slot, about 14x less than NVLink. For tensor-parallel inference or distributed training across cards, this is the wrong hardware and no amount of cost advantage fixes it. 70B and larger models on a single card force quantization or multi-GPU setups that change the economics entirely.
There is one market signal worth noting, because it runs against the pattern for every other previous-generation part. As Blackwell volume ramps, A100 and H100 hardware is flowing into secondary channels and pricing is softening. The L40S is holding most of its value. Demand for the current inference favorite has not moved, which tells you something about where real serving workloads are actually running rather than where benchmark attention goes.
Get a quote for the tier that matches the workload.
Not the one that tops the charts. No buyer fees. For single GPUs, packet.ai handles self-serve access with 24/7 human support.
Because the price gap is larger than the throughput gap. L40S runs roughly 5.33x cheaper per hour while the H100's LLM serving advantage runs 2.7x to 5x. Fewer tokens per second, more tokens per dollar. Measured cost lands at $0.023 per million against the H100's $0.026.
Model size decides it. L40S for under 30B, where its FP8 support matters and the bandwidth deficit does not bite. A100 for 30B to 70B at FP8, where 80GB and 2 TB/s of bandwidth do. The A100 has no FP8 support at all, which is the difference most comparisons bury.
No. The L40S has no NVLink and uses PCIe Gen4 only. PCIe 4.0 delivers roughly 64 GB/s bidirectional on a sixteen-lane slot, about 14x less than NVLink. For tensor-parallel work you need H100 SXM or B200 SXM.
It is the strongest case for the card. 48GB handles SDXL at any standard resolution and Flux.1 up to 1080p, generating at a rate comparable to H100 while costing 5.33x less per hour. Most image deployments are latency-focused rather than batch-throughput-focused, which is where the H100 advantage would compound.
Demand for it as an inference card has not moved as Blackwell ramps. A100 and H100 hardware is flowing into secondary channels from fleet refreshes and pricing is drifting down. The L40S is not part of that flow at the same rate, which is a signal about where production serving workloads actually sit.
Last reviewed: 27 August 2026. Cost-per-token and throughput figures from Cyfuture's L40S vs A100 vs H100 comparison and Spheron's L40S vs H100 decision guide. Specification data from NVIDIA L40S and H100 datasheets via GetDeploying. Image generation and embedding workload analysis from Spheron. Pricing from GetDeploying's L40S provider index and packet.ai L40S pricing, July 2026. Secondary market observations from PCSP's used GPU server buying guide, July 2026. Browse current GPU cluster availability on GPUaaS.com.