How much VRAM does BLOOM 176B z need?
BLOOM 176B z needs about 105.1 GB of VRAM at Q4_K_M, or 183.6 GB at Q8_0 and 345.0 GB at FP16. No consumer card holds it at Q4_K_M. The smallest fit in our set is the H200 141GB (141 GB). Weights plus runtime overhead; with only 2K of context the cache stays small.
Weights-plus-runtime footprint across the six most common quantizations. KV cache is context-dependent and comes on top. The 2K window is short, so weights dominate the budget; size it precisely in the GPU Memory Calculator.
Can your GPU run BLOOM 176B z?
Pick your card — the answer below is computed by the same engine that produces every figure on this page.
H200 141GB runs BLOOM 176B z at Q4_K_M with 28.9 GB to spare; comfortable at this size, where most cards cannot hold the weights at all.
Weights plus runtime overhead. KV cache is not included — it depends on context length and this model’s published architecture.
| Weights (Q4_K_M) | 99.4 GB |
| Runtime overhead | 5.7 GB |
| Estimated total | 105.1 GB |
| Usable VRAM (95% of 141 GB) | 133.9 GB |
Across every GPU we track
- RTX 3060 12GBNot enough-93.7 GB
- RTX 4060 Ti 16GBNot enough-89.9 GB
- RTX 4070 Ti SUPERNot enough-89.9 GB
- RTX 3090Not enough-82.3 GB
- RTX 4090Not enough-82.3 GB
- RTX 5090Not enough-74.7 GB
- A100 40GBNot enough-67.1 GB
- RTX A6000Not enough-59.5 GB
- L40SNot enough-59.5 GB
- A100 80GBNot enough-29.1 GB
- H100 80GBNot enough-29.1 GB
- H200 141GBRecommended+28.9 GB
Estimates, not guarantees: real usage moves with runtime, driver, batch size and context. Size a specific context window in the GPU Memory Calculator.
How was this number calculated? Every figure here comes from Bitpute’s documented calculation methodology — parameters, precision, quantization, runtime overhead and usable VRAM, each formula written out in full.
Memory by quantization
| Quant | Weights | + runtime overhead |
|---|---|---|
| FP16 | 327.8 GB | 345.0 GB |
| Q8_0 | 174.2 GB | 183.6 GB |
| Q6_K | 134.4 GB | 141.9 GB |
| Q5_K_M | 116.6 GB | 123.2 GB |
| Q4_K_M | 99.4 GB | 105.1 GB |
| Q4_0 | 92.2 GB | 97.6 GB |
Overhead = 0.75 GB + 5% of weights (CUDA context, buffers). Add KV cache on top. With only 2K of context, weights dominate the budget.
Which quantization should you actually run?
For BLOOM 176B z, the single-GPU entry point is Q4_K_M (105.1 GB) on H200 141GB, with plenty of headroom for long context. Higher quality (Q6_K 141.9 GB, FP16 345.0 GB) needs a 48 GB+ card or multiple GPUs — for occasional use, renting is usually cheaper than buying.
Single-GPU fit at Q4_K_M (105.1 GB + KV)
RTX 3060 12GBRTX 4060 Ti 16GBRTX 4070 Ti SUPERRTX 3090RTX 4090RTX 5090A100 40GBRTX A6000L40SA100 80GBH100 80GBH200 141GB
Where the memory goes
Weights and overhead are exact for Q4_K_M. The KV bar is a generic 32-layer transformer at 8K; at this scale the weights dominate, but long context still adds gigabytes — size it in the calculator is for.
Common questions
How much VRAM does BLOOM 176B z need?
BLOOM 176B z needs about 105.1 GB of VRAM at Q4_K_M, 183.6 GB at Q8_0, or 345.0 GB at FP16. That is weights plus about 0.75 GB of runtime overhead and 5% of weight size; KV cache is additional and depends on context length.
What GPU can run BLOOM 176B z?
At Q4_K_M the smallest card in our set that fits is the H200 141GB (141 GB usable at 95%). At FP16 no single GPU in our set is sufficient.
Can BLOOM 176B z run on 24 GB?
Not at Q4_K_M: it needs about 105.1 GB, more than the ~22.8 GB usable on a 24 GB card. Use a larger card, multiple GPUs, or offload.
Same memory footprint as 1 other checkpoint
BLOOM 176B z has an identical parameter count to this, so every figure on this page applies to it unchanged. What varies is different checkpoint roles (base) — not the size. Choose on capability and licence; the memory budget is identical.
More in the BLOOM family
BLOOM 560MBLOOM 560M zBLOOM 1.7BBLOOM 1.7B zBLOOM 3BBLOOM 3B z