GPU guide
H100 80GB for AI: what it actually runs
In short
- Holds 473 of 506 catalogue models at Q4_K_M on 80 GB of VRAM.
- Largest fit: Pixtral Large 124B Instruct at 74.3 GB.
- Memory bandwidth 3350 GB/s — rank 2 of 12 in our set. Decode speed tracks this figure once a model fits.
- 700 W board power — about 168 kWh a month at 8 hours a day. Apply your own tariff.
Form factor: SXM5. The figures on this page are for the H100 SXM5 module: 3.35 TB/s and 700 W. The PCIe card is ~2 TB/s at 350 W — roughly 40% less bandwidth, which materially lowers decode speed. Source: NVIDIA datasheet.
The current default for production inference. Same 80 GB as the A100 80GB, but 3350 GB/s of bandwidth — a two-thirds jump that translates almost directly into tokens per second on memory-bound LLM work.
AI suitability
Everything the A100 80GB fits, the H100 fits and serves markedly faster; the theoretical ceiling on Llama 3.3 70B at Q4 is around 78 tokens per second. Its 700 W draw is a rack consideration, not a home one — this is a card you rent by the hour far more often than you own.
For LLM inference the number that matters most is memory bandwidth, because generating each token means reading the whole model. On the H100 80GB that puts the theoretical ceiling around 688 tokens/second on Llama 3.1 8B at Q4_K_M, and about 76 tokens/second on the biggest model it comfortably holds, Qwen2.5 72B — real-world throughput lands below these ceilings. At Q4_K_M with an 8K context this card comfortably fits 12 of the 12 models in our database; at the near-lossless Q8_0, 10.
Which models fit the H100 80GB?
Computed at an 8K context (or the model's own cap). ✓ fits comfortably (≤95% of 80 GB) · ~ tight · ✗ doesn't fit. Every model links to its own guide.
| Model | Q4_K_M | Q8_0 |
|---|---|---|
| Llama 3.2 1B | 1.7 GB ✓ | 2.3 GB ✓ |
| Llama 3.2 3B | 3.5 GB ✓ | 5.0 GB ✓ |
| Mistral 7B | 6.0 GB ✓ | 9.3 GB ✓ |
| Qwen2.5 7B | 5.7 GB ✓ | 9.1 GB ✓ |
| Llama 3.1 8B | 6.5 GB ✓ | 10.1 GB ✓ |
| Gemma 2 9B | 8.9 GB ✓ | 13.0 GB ✓ |
| Qwen2.5 14B | 11.0 GB ✓ | 17.6 GB ✓ |
| Gemma 2 27B | 19.8 GB ✓ | 31.9 GB ✓ |
| Qwen2.5 32B | 22.2 GB ✓ | 36.8 GB ✓ |
| DeepSeek-R1 32B | 22.2 GB ✓ | 36.8 GB ✓ |
| Llama 3.3 70B | 45.1 GB ✓ | 76.6 GB ~ |
| Qwen2.5 72B | 46.3 GB ✓ | 78.8 GB ~ |
Electricity
The board is rated at 700 W. Run it under sustained load 8 hours a day and that's about 168 kWh a month — roughly $20/month at an example rate of $0.12/kWh (set your own tariff in the workspace). Idle and light chat draw far less; the figure above is the worst case, not the typical bill.
Renting instead of buying
This is a cloud-native card — the H100 80GB is something you rent far more often than you buy. Hourly rates move constantly, so we don't print them here; the workspace holds the current figures and shows rental cost against electricity for any model you pick.
Alternatives
48 GB · 864 GB/s · 350 W
Step down — keeps most small models, saves money and watts.
See this card →80 GB · 2039 GB/s · 400 W
Same memory tier — the fit list is identical; bandwidth and price decide.
See this card →141 GB · 4800 GB/s · 700 W
Step up — the next memory tier and what it unlocks.
See this card →Compare any two of these head-to-head — speed on the same model, cost, power — in GPU Compare.
H100 80GB in the Bitpute graph
Everything this card connects to — models it runs, where to rent it, when we recommend it, and its nearest rivals. Derived from the same data as the calculators; reviewed July 2026.
Runs these models
Largest fits at Q4_K_M (of 473 that fit):
Pixtral Large 124B InstructMistral Large 123B (2407) InstructGPT-OSS 120B (5.1B active) InstructCommand A 111B ChatQwen 1.5 110BEven at FP16 (417 fit):
Seed-OSS 36B InstructAya 23 35B ChatCommand R 35B (v01) ChatBrowse all 506 model profiles →Weights + overhead only; KV cache comes on top.
Rent it in the cloud
17 providers stock this class of card:
Vast.aiTogether AINebiusHyperstack (NexGen Cloud)RunPodDigitalOcean Gradient (Paperspace)TensorDockLambdaTypically cheapest tier: Vast.ai. On-demand this card runs about $2.85/hr.
Compare all 18 providers →When we recommend it
This card never tops an award category — a neighbour beats it on price, bandwidth or power at every requirement size. It can still be the right buy at the right street price.
Run your own numbers →Why bandwidth matters more than core count for token generation: memory bandwidth explained.
Capacity against speed
The two axes are independent. Position on the horizontal decides which models fit; position on the vertical decides how fast they decode once they do.
If a model will not load on this card
In rough order of likelihood. For reference, the largest model in our catalogue that fits this card at Q4_K_M is Pixtral Large 124B Instruct at 74.3 GB.
- The model is larger than usable VRAM, not nameplate VRAM. This card reports 80 GB but roughly 76.0 GB is available to a model after the display buffer, driver and CUDA context. A model sized against the nameplate figure will appear to fit and then fail.
- KV cache grew past the headroom. Weights are fixed; the cache is not. A load that succeeds at 2K context can fail at 32K on the same card, because the cache comes out of the same budget. If it loaded yesterday and fails today, context length is the first thing to check. See KV cache.
- The quantization is heavier than assumed. Q4_K_M is roughly 4.85 bits per weight, not 4. On a large model that difference is gigabytes. Confirm which file you actually downloaded — see which quantization to run.
- Something else is already holding VRAM. A browser with hardware acceleration, another model still resident, or a previous process that did not release memory. On Linux
nvidia-smishows what is allocated. - It loaded but generation is very slow. That usually means layers were offloaded to system RAM rather than the load failing outright. The model runs, but every offloaded layer crosses PCIe on each token. Reduce context, drop a quantization level, or use a card with more VRAM.
Common questions
Is the H100 80GB good for AI and local LLMs?
Yes, within its tier: with 80 GB of VRAM and 3350 GB/s of memory bandwidth it comfortably runs 12 of the 12 models in our database at Q4_K_M with an 8K context, the largest being Qwen2.5 72B. Memory bandwidth caps generation speed at roughly 688 tokens/second on an 8B model in theory, with real-world results lower.
What is the biggest model an H100 80GB can run?
Qwen2.5 72B at Q4_K_M with an 8K context is the largest comfortable fit (46.3 GB of 80 GB). Larger models need a bigger card or a multi-GPU split.
How much electricity does an H100 80GB use?
The board is rated at 700 W. Under sustained load for 8 hours a day that is about 168 kWh a month — roughly $20/month at an example rate of $0.12/kWh. Idle draw is far lower, so light interactive use costs much less.
Data centre cards, ranked by VRAM
Server GPUs: HBM bandwidth, NVLink, MIG partitioning, rack thermals. Every card in this tier, smallest memory first — the one you are reading is highlighted. Cards shown without a memory figure are not in our calculation engine yet.
Where it sits in the stack
The H100 80GB is a Hopper-generation part with HBM3 memory and 4th-generation Tensor Cores. Memory type is why the bandwidth figure above lands where it does: HBM stacks feed far more bandwidth per watt than GDDR, and decode speed on a model that already fits is set by bandwidth, not by core count.
For multi-GPU work it supports NVLink — 900 GB/s between GPUs — which matters when a model is split across cards, because tensor-parallel inference moves activations between GPUs on every token.
Its Tensor Cores handle FP16, BF16, INT8 and FP8. FP8 halves activation memory against FP16 where a runtime supports it; Ampere cards do not have it. Most local inference runs weights at 4-6 bits via GGUF quantization rather than a native Tensor Core format, so these matter most for training and for server runtimes such as TensorRT-LLM and vLLM.
Hopper adds the Transformer Engine and FP8, which cuts activation memory and speeds up attention on supported runtimes. MIG can slice it into as many as seven isolated instances, so one board can serve several small models at once. None of this changes the VRAM arithmetic on this page — weights and KV cache still have to fit.
On the software side it is a CUDA device like any other NVIDIA card, so PyTorch, TensorRT, vLLM, llama.cpp and Ollama all run on it unchanged. What differs between cards is not compatibility but how much fits and how fast it decodes — which is what the numbers above measure. For the quantization formats these runtimes expect, and how much room the GPU Memory Calculator says you have left, start there.
Head to head
Pre-computed comparisons against the cards it is usually weighed against:
A100 80GB vs H100 80GBH100 80GB vs H200 141GBRTX 4090 vs H100 80GBRTX 5090 vs H100 80GB