Skip to content
Bitpute

GPU guide

A100 40GB for AI: what it actually runs

By Bitpute · Published 12 July 2026 · Updated 24 July 2026 · Specs verified against manufacturer datasheets 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error · NVIDIA A100 datasheet

In short

40 GB VRAM 1555 GB/s bandwidth 400 W TDP datacenter

Form factor: SXM4. The figures on this page are for the A100 40GB SXM4 module: 1,555 GB/s and 400 W. The PCIe card has the same 1,555 GB/s bandwidth but a 250 W TDP, so power and running-cost figures differ. Source: NVIDIA datasheet.

The older cloud staple — and a lesson in reading the memory number. The A100 40GB has datacenter-grade bandwidth (1555 GB/s) yet its 40 GB falls about 5 GB short of Llama 3.3 70B at Q4_K_M, a model the slower 48 GB cards handle.

AI suitability

So its real role today is speed on the mid-sizes: the 32B class at Q4 or even Q8 runs with room to spare and generates quickly. If your target is 70B on a single card, skip this tier entirely; if your target is fast 32B inference on rented hardware, it is often the value pick.

For LLM inference the number that matters most is memory bandwidth, because generating each token means reading the whole model. On the A100 40GB that puts the theoretical ceiling around 319 tokens/second on Llama 3.1 8B at Q4_K_M, and about 78 tokens/second on the biggest model it comfortably holds, Qwen2.5 32B — real-world throughput lands below these ceilings. At Q4_K_M with an 8K context this card comfortably fits 10 of the 12 models in our database; at the near-lossless Q8_0, 10.

Which models fit the A100 40GB?

Computed at an 8K context (or the model's own cap). ✓ fits comfortably (≤95% of 40 GB) · ~ tight · ✗ doesn't fit. Every model links to its own guide.

ModelQ4_K_MQ8_0
Llama 3.2 1B1.7 GB 2.3 GB
Llama 3.2 3B3.5 GB 5.0 GB
Mistral 7B6.0 GB 9.3 GB
Qwen2.5 7B5.7 GB 9.1 GB
Llama 3.1 8B6.5 GB 10.1 GB
Gemma 2 9B8.9 GB 13.0 GB
Qwen2.5 14B11.0 GB 17.6 GB
Gemma 2 27B19.8 GB 31.9 GB
Qwen2.5 32B22.2 GB 36.8 GB
DeepSeek-R1 32B22.2 GB 36.8 GB
Llama 3.3 70B45.1 GB 76.6 GB
Qwen2.5 72B46.3 GB 78.8 GB

Electricity

The board is rated at 400 W. Run it under sustained load 8 hours a day and that's about 96 kWh a month — roughly $12/month at an example rate of $0.12/kWh (set your own tariff in the workspace). Idle and light chat draw far less; the figure above is the worst case, not the typical bill.

Renting instead of buying

This is a cloud-native card — the A100 40GB is something you rent far more often than you buy. Hourly rates move constantly, so we don't print them here; the workspace holds the current figures and shows rental cost against electricity for any model you pick.

Alternatives

RTX 5090

32 GB · 1792 GB/s · 575 W

Step down — keeps most small models, saves money and watts.

See this card →
L40S

48 GB · 864 GB/s · 350 W

Step up — the next memory tier and what it unlocks.

See this card →

Compare any two of these head-to-head — speed on the same model, cost, power — in GPU Compare.

← L40S · All GPUs · A100 80GB →

A100 40GB in the Bitpute graph

Everything this card connects to — models it runs, where to rent it, when we recommend it, and its nearest rivals. Derived from the same data as the calculators; reviewed July 2026.

Rent it in the cloud

1 providers stock this class of card:

Vast.ai

Typically cheapest tier: Vast.ai. On-demand this card runs about $0.85/hr.

Compare all 18 providers →

When we recommend it

Best value30.5–38 GB required
Best for fine-tuning20.5–25 GB required

Ranges where our recommendation engine picks this card, given the VRAM a workload needs.

Run your own numbers →

Why bandwidth matters more than core count for token generation: memory bandwidth explained.

Capacity against speed

The two axes are independent. Position on the horizontal decides which models fit; position on the vertical decides how fast they decode once they do. Cards up and to the left of it hold less but read faster; RTX A6000 sits lower-right — more memory, less bandwidth.

VRAM against memory bandwidth, A100 40GB highlightedScatter plot of the 12 GPUs in our set, VRAM on the horizontal axis against memory bandwidth on the vertical. A100 40GB sits at 40 GB and 1555 GB/s.01296259238885184VRAM (GB) →Bandwidth (GB/s) →122448141A100 40GB

If a model will not load on this card

In rough order of likelihood. For reference, the largest model in our catalogue that fits this card at Q4_K_M is Qwen 2 57B-A14B Instruct at 34.5 GB.

  1. The model is larger than usable VRAM, not nameplate VRAM. This card reports 40 GB but roughly 38.0 GB is available to a model after the display buffer, driver and CUDA context. A model sized against the nameplate figure will appear to fit and then fail.
  2. KV cache grew past the headroom. Weights are fixed; the cache is not. A load that succeeds at 2K context can fail at 32K on the same card, because the cache comes out of the same budget. If it loaded yesterday and fails today, context length is the first thing to check. See KV cache.
  3. The quantization is heavier than assumed. Q4_K_M is roughly 4.85 bits per weight, not 4. On a large model that difference is gigabytes. Confirm which file you actually downloaded — see which quantization to run.
  4. Something else is already holding VRAM. A browser with hardware acceleration, another model still resident, or a previous process that did not release memory. On Linux nvidia-smi shows what is allocated.
  5. It loaded but generation is very slow. That usually means layers were offloaded to system RAM rather than the load failing outright. The model runs, but every offloaded layer crosses PCIe on each token. Reduce context, drop a quantization level, or use a card with more VRAM.

Common questions

Is the A100 40GB good for AI and local LLMs?

Yes, within its tier: with 40 GB of VRAM and 1555 GB/s of memory bandwidth it comfortably runs 10 of the 12 models in our database at Q4_K_M with an 8K context, the largest being Qwen2.5 32B. Memory bandwidth caps generation speed at roughly 319 tokens/second on an 8B model in theory, with real-world results lower.

What is the biggest model an A100 40GB can run?

Qwen2.5 32B at Q4_K_M with an 8K context is the largest comfortable fit (22.2 GB of 40 GB). Larger models need a bigger card or a multi-GPU split.

How much electricity does an A100 40GB use?

The board is rated at 400 W. Under sustained load for 8 hours a day that is about 96 kWh a month — roughly $12/month at an example rate of $0.12/kWh. Idle draw is far lower, so light interactive use costs much less.

Data centre cards, ranked by VRAM

Server GPUs: HBM bandwidth, NVLink, MIG partitioning, rack thermals. Every card in this tier, smallest memory first — the one you are reading is highlighted. Cards shown without a memory figure are not in our calculation engine yet.

A100 40GB40 GBL40S48 GBA100 80GB80 GBH100 80GB80 GBH200 141GB141 GB

Where it sits in the stack

The A100 40GB is an Ampere-generation part with HBM2 memory and 3rd-generation Tensor Cores. Memory type is why the bandwidth figure above lands where it does: HBM stacks feed far more bandwidth per watt than GDDR, and decode speed on a model that already fits is set by bandwidth, not by core count.

For multi-GPU work it supports NVLink — 600 GB/s between GPUs — which matters when a model is split across cards, because tensor-parallel inference moves activations between GPUs on every token.

Its Tensor Cores handle FP16, BF16 and INT8. There is no FP8 path here — that arrives with Ada and Hopper — so the practical low-precision options are INT8 and the GGUF integer quants. Most local inference runs weights at 4-6 bits via GGUF quantization rather than a native Tensor Core format, so these matter most for training and for server runtimes such as TensorRT-LLM and vLLM.

MIG can slice it into as many as seven isolated instances, so one board can serve several small models at once. None of this changes the VRAM arithmetic on this page — weights and KV cache still have to fit.

On the software side it is a CUDA device like any other NVIDIA card, so PyTorch, TensorRT, vLLM, llama.cpp and Ollama all run on it unchanged. What differs between cards is not compatibility but how much fits and how fast it decodes — which is what the numbers above measure. For the quantization formats these runtimes expect, and how much room the GPU Memory Calculator says you have left, start there.

Head to head

Pre-computed comparisons against the cards it is usually weighed against:

A100 40GB vs A100 80GBA100 40GB vs L40S