How much VRAM does DeepSeek-V2 236B (21B active) need?
DeepSeek-V2 236B (21B active) needs about 140.7 GB of VRAM at Q4_K_M, or 246.0 GB at Q8_0 and 462.3 GB at FP16. No single GPU in our set holds it even at Q4_K_M; it needs multi-GPU or offload. Weights plus runtime overhead. Budget separately for KV cache, which at 128K is a material share.
Weights-plus-runtime footprint across the six most common quantizations. KV cache is context-dependent and comes on top. At 128K it is gigabytes on its own, before the weights; size it precisely in the GPU Memory Calculator.
This is a mixture-of-experts model: all 236B parameters must sit in memory, but only about 21B are active per token — so it loads like a 236B model and runs closer to a 21B one.
This is the base checkpoint: pretrained on raw text and not tuned to follow instructions, so it continues text rather than answering prompts. That makes it the right starting point if you plan to fine-tune your own model. For chat or assistant work, use DeepSeek-V2 236B (21B active) Chat — same parameter count, so the memory figures below are identical.
Can your GPU run DeepSeek-V2 236B (21B active)?
Pick your card — the answer below is computed by the same engine that produces every figure on this page.
H200 141GB cannot hold DeepSeek-V2 236B (21B active) at Q4_K_M: it needs 140.7 GB against 133.9 GB usable; a heavier quantization, a larger card, or splitting across GPUs are the ways forward; as a mixture-of-experts model all 236B parameters must be resident even though only 21B are active per token.
Weights plus runtime overhead. KV cache is not included — it depends on context length and this model’s published architecture.
| Weights (Q4_K_M) | 133.2 GB |
| Runtime overhead | 7.4 GB |
| Estimated total | 140.7 GB |
| Usable VRAM (95% of 141 GB) | 133.9 GB |
Across every GPU we track
- RTX 3060 12GBNot enough-129.3 GB
- RTX 4060 Ti 16GBNot enough-125.5 GB
- RTX 4070 Ti SUPERNot enough-125.5 GB
- RTX 3090Not enough-117.9 GB
- RTX 4090Not enough-117.9 GB
- RTX 5090Not enough-110.3 GB
- A100 40GBNot enough-102.7 GB
- RTX A6000Not enough-95.1 GB
- L40SNot enough-95.1 GB
- A100 80GBNot enough-64.7 GB
- H100 80GBNot enough-64.7 GB
- H200 141GBNot enough-6.7 GB
Estimates, not guarantees: real usage moves with runtime, driver, batch size and context. Size a specific context window in the GPU Memory Calculator.
How was this number calculated? Every figure here comes from Bitpute’s documented calculation methodology — parameters, precision, quantization, runtime overhead and usable VRAM, each formula written out in full.
Memory by quantization
| Quant | Weights | + runtime overhead |
|---|---|---|
| FP16 | 439.6 GB | 462.3 GB |
| Q8_0 | 233.5 GB | 246.0 GB |
| Q6_K | 180.2 GB | 190.0 GB |
| Q5_K_M | 156.3 GB | 164.9 GB |
| Q4_K_M | 133.2 GB | 140.7 GB |
| Q4_0 | 123.6 GB | 130.6 GB |
Overhead = 0.75 GB + 5% of weights (CUDA context, buffers). Add KV cache on top. Weights dominate at this size, but a 128K window still costs gigabytes.
Which quantization should you actually run?
DeepSeek-V2 236B (21B active) is a multi-GPU or datacenter model: even Q4_K_M needs about 140.7 GB, beyond any single card. Plan for several 80 GB GPUs or a hosted endpoint; FP16 (462.3 GB) is server-only. As a mixture-of-experts model it must hold all 236B in VRAM but activates only about 21B per token, so it runs much faster than its footprint suggests — memory is the limit here, not speed.
Single-GPU fit at Q4_K_M (140.7 GB + KV)
RTX 3060 12GBRTX 4060 Ti 16GBRTX 4070 Ti SUPERRTX 3090RTX 4090RTX 5090A100 40GBRTX A6000L40SA100 80GBH100 80GBH200 141GB
Common questions
How much VRAM does DeepSeek-V2 236B (21B active) need?
DeepSeek-V2 236B (21B active) needs about 140.7 GB of VRAM at Q4_K_M, 246.0 GB at Q8_0, or 462.3 GB at FP16. That is weights plus about 0.75 GB of runtime overhead and 5% of weight size; KV cache is additional and depends on context length.
What GPU can run DeepSeek-V2 236B (21B active)?
No single GPU in our set fits it at Q4_K_M; it needs multiple GPUs or CPU offload. At FP16 no single GPU in our set is sufficient.
Can DeepSeek-V2 236B (21B active) run on 24 GB?
Not at Q4_K_M: it needs about 140.7 GB, more than the ~22.8 GB usable on a 24 GB card. Use a larger card, multiple GPUs, or offload.
Same memory footprint as 1 other checkpoint
DeepSeek-V2 236B (21B active) has an identical parameter count to this, so every figure on this page applies to it unchanged. What varies is different checkpoint roles (chat-tuned) — not the size. Choose on capability and licence; the memory budget is identical.
DeepSeek-V2 236B (21B active) Chat
More in the DeepSeek-V2 family
DeepSeek-V2 Lite 16B (2.4B active)DeepSeek-V2 Lite 16B (2.4B active) ChatDeepSeek-V2 236B (21B active) Chat