A100 40GB vs L40S
The L40S holds 33 models from our catalog that the A100 40GB cannot fit at Q4_K_M, and decodes about 44% slower on memory bandwidth (864 GB/s against 1555 GB/s). Decode speed on a memory-bound workload scales with bandwidth, so that ratio is the number that matters once the model fits at all.
Specs side by side
| Spec | A100 40GB | L40S |
|---|---|---|
| VRAM | 40 GB | 48 GB |
| Memory bandwidth | 1555 GB/s | 864 GB/s |
| TDP | 400 W | 350 W |
| Models that fit (Q4_K_M) | 427 of 506 | 460 of 506 |
Form factor: A100 40GB figures are for the SXM4 module. PCIe/NVL variants differ in bandwidth and TDP — see the A100 40GB page and methodology.
What the extra VRAM actually buys
33 models in our catalog fit on the L40S at Q4_K_M but not on the A100 40GB. The largest of them:
Qwen 2 VL 72B InstructQwen 2.5 VL 72B InstructMolmo 72BQwen 2 72BQwen 2 72B InstructQwen 2.5 72B
Largest model each can run
At Q4_K_M the biggest fit in our catalog is Qwen 2 57B-A14B Instruct (34.5 GB) on the A100 40GB, and Qwen 2 VL 72B Instruct (44.3 GB) on the L40S. Both figures are weights plus runtime overhead; KV cache for your context is extra.
Power and running cost
At 8 hours a day the A100 40GB draws about 96 kWh a month against 84 kWh for the L40S — a difference of 12 kWh. Multiply by your own tariff; we do not assume one. Board TDP is the rated figure, not measured draw.
FAQ
A100 40GB vs L40S: which should I buy for local AI?
The L40S fits 460 of the 506 models in our catalog at Q4_K_M against 427 for the A100 40GB, and has about 0.6× the memory bandwidth, which is what sets decode speed once a model fits. If the models you run already fit the A100 40GB, the extra spend buys context length and speed, not a new model tier.
Is the L40S faster than the A100 40GB?
No — despite the larger VRAM it has lower memory bandwidth (864 GB/s against 1555 GB/s), so on memory-bound decoding it is roughly 44% slower. It holds bigger models; it does not run them faster. This is an upper bound from bandwidth alone, not a measured benchmark.
How much more VRAM does the L40S have?
8 GB more (48 GB vs 40 GB). That difference lets it hold 33 additional models from our catalog at Q4_K_M.
Need a third card in the mix? The interactive compare tool lines up any three. To size KV cache for your context length, use the GPU Memory Calculator.
Related matchups
Other head-to-heads involving these cards:
RTX A6000 vs L40SA100 40GB vs A100 80GBRTX 5090 vs L40SRTX 5090 vs A100 40GBA100 40GB vs RTX A6000
New to sizing? What VRAM is and why bandwidth sets decode speed cover the basics, and quantization explains the formats in the table above. Common mistakes when sizing VRAM lists the traps.