RTX 5090 vs A100 40GB
The A100 40GB holds 4 models from our catalog that the RTX 5090 cannot fit at Q4_K_M, and decodes about 13% slower on memory bandwidth (1555 GB/s against 1792 GB/s). Decode speed on a memory-bound workload scales with bandwidth, so that ratio is the number that matters once the model fits at all.
Specs side by side
| Spec | RTX 5090 | A100 40GB |
|---|---|---|
| VRAM | 32 GB | 40 GB |
| Memory bandwidth | 1792 GB/s | 1555 GB/s |
| TDP | 575 W | 400 W |
| Models that fit (Q4_K_M) | 423 of 506 | 427 of 506 |
Form factor: A100 40GB figures are for the SXM4 module. PCIe/NVL variants differ in bandwidth and TDP — see the A100 40GB page and methodology.
What the extra VRAM actually buys
4 models in our catalog fit on the A100 40GB at Q4_K_M but not on the RTX 5090. The largest of them:
Qwen 2 57B-A14B InstructJamba 1.5 Mini 52B InstructJamba v0.1 52B (12B active)Nemotron Llama-3.1 51B Instruct
Largest model each can run
At Q4_K_M the biggest fit in our catalog is Nemotron Super 49B Instruct (29.8 GB) on the RTX 5090, and Qwen 2 57B-A14B Instruct (34.5 GB) on the A100 40GB. Both figures are weights plus runtime overhead; KV cache for your context is extra.
Power and running cost
At 8 hours a day the RTX 5090 draws about 138 kWh a month against 96 kWh for the A100 40GB — a difference of 42 kWh. Multiply by your own tariff; we do not assume one. Board TDP is the rated figure, not measured draw.
FAQ
RTX 5090 vs A100 40GB: which should I buy for local AI?
The A100 40GB fits 427 of the 506 models in our catalog at Q4_K_M against 423 for the RTX 5090, and has about 0.9× the memory bandwidth, which is what sets decode speed once a model fits. If the models you run already fit the RTX 5090, the extra spend buys context length and speed, not a new model tier.
Is the A100 40GB faster than the RTX 5090?
No — despite the larger VRAM it has lower memory bandwidth (1555 GB/s against 1792 GB/s), so on memory-bound decoding it is roughly 13% slower. It holds bigger models; it does not run them faster. This is an upper bound from bandwidth alone, not a measured benchmark.
How much more VRAM does the A100 40GB have?
8 GB more (40 GB vs 32 GB). That difference lets it hold 4 additional models from our catalog at Q4_K_M.
Need a third card in the mix? The interactive compare tool lines up any three. To size KV cache for your context length, use the GPU Memory Calculator.
Related matchups
Other head-to-heads involving these cards:
RTX 4090 vs RTX 5090RTX 3090 vs RTX 5090RTX 5090 vs RTX A6000A100 40GB vs L40SA100 40GB vs A100 80GB
New to sizing? What VRAM is and why bandwidth sets decode speed cover the basics, and quantization explains the formats in the table above. Common mistakes when sizing VRAM lists the traps.