Skip to content
Bitpute

How much VRAM does DeepSeek-V3 0324 671B Chat need?

By Bitpute · Published 12 July 2026 · Updated 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error · Figures re-checked against the calculation engine on every build

DeepSeek-V3 0324 671B Chat needs about 398.5 GB of VRAM at Q4_K_M, or 697.9 GB at Q8_0 and 1313.1 GB at FP16. No single GPU in our set holds it even at Q4_K_M; it needs multi-GPU or offload. Weights plus runtime overhead. Budget separately for KV cache, which at 128K is a material share.

Weights-plus-runtime footprint across the six most common quantizations. KV cache is context-dependent and comes on top. At 128K it is gigabytes on its own, before the weights; size it precisely in the GPU Memory Calculator.

This is a mixture-of-experts model: all 671B parameters must sit in memory, but only about 37B are active per token — so it loads like a 671B model and runs closer to a 37B one.

This is the instruction-tuned checkpoint: the same architecture and the same memory footprint as the base model, fine-tuned to follow prompts and hold a conversation. If you intend to fine-tune on your own data, start from DeepSeek-V3 671B (37B active) instead.

Parameters671B
Active / token37B
Context128K
VendorDeepSeek
LicenseMIT
Released2024

Can your GPU run DeepSeek-V3 0324 671B Chat?

Pick your card — the answer below is computed by the same engine that produces every figure on this page.

No — not enough VRAM
Estimated VRAM398.5 GB
GPU VRAM141 GB
Short by264.6 GB

H200 141GB cannot hold DeepSeek-V3 0324 671B Chat at Q4_K_M: it needs 398.5 GB against 133.9 GB usable; a heavier quantization, a larger card, or splitting across GPUs are the ways forward; as a mixture-of-experts model all 671B parameters must be resident even though only 37B are active per token.

Weights plus runtime overhead. KV cache is not included — it depends on context length and this model’s published architecture.

Why
Weights (Q4_K_M)378.9 GB
Runtime overhead19.7 GB
Estimated total398.5 GB
Usable VRAM (95% of 141 GB)133.9 GB

Across every GPU we track

Estimates, not guarantees: real usage moves with runtime, driver, batch size and context. Size a specific context window in the GPU Memory Calculator.

How was this number calculated? Every figure here comes from Bitpute’s documented calculation methodology — parameters, precision, quantization, runtime overhead and usable VRAM, each formula written out in full.

Memory by quantization

QuantWeights+ runtime overhead
FP161249.8 GB1313.1 GB
Q8_0664.0 GB697.9 GB
Q6_K512.4 GB538.8 GB
Q5_K_M444.5 GB467.4 GB
Q4_K_M378.9 GB398.5 GB
Q4_0351.5 GB369.8 GB

Overhead = 0.75 GB + 5% of weights (CUDA context, buffers). Add KV cache on top. Weights dominate at this size, but a 128K window still costs gigabytes.

Which quantization should you actually run?

DeepSeek-V3 0324 671B Chat is a multi-GPU or datacenter model: even Q4_K_M needs about 398.5 GB, beyond any single card. Plan for several 80 GB GPUs or a hosted endpoint; FP16 (1313.1 GB) is server-only. As a mixture-of-experts model it must hold all 671B in VRAM but activates only about 37B per token, so it runs much faster than its footprint suggests — memory is the limit here, not speed.

Single-GPU fit at Q4_K_M (398.5 GB + KV)

RTX 3060 12GBRTX 4060 Ti 16GBRTX 4070 Ti SUPERRTX 3090RTX 4090RTX 5090A100 40GBRTX A6000L40SA100 80GBH100 80GBH200 141GB

Size it exactly →Rent a GPU for it →

Common questions

How much VRAM does DeepSeek-V3 0324 671B Chat need?

DeepSeek-V3 0324 671B Chat needs about 398.5 GB of VRAM at Q4_K_M, 697.9 GB at Q8_0, or 1313.1 GB at FP16. That is weights plus about 0.75 GB of runtime overhead and 5% of weight size; KV cache is additional and depends on context length.

What GPU can run DeepSeek-V3 0324 671B Chat?

No single GPU in our set fits it at Q4_K_M; it needs multiple GPUs or CPU offload. At FP16 no single GPU in our set is sufficient.

Can DeepSeek-V3 0324 671B Chat run on 24 GB?

Not at Q4_K_M: it needs about 398.5 GB, more than the ~22.8 GB usable on a 24 GB card. Use a larger card, multiple GPUs, or offload.

Same memory footprint as 1 other checkpoint

DeepSeek-V3 0324 671B Chat has an identical parameter count to this, so every figure on this page applies to it unchanged. What varies is different checkpoint roles (base) — not the size. Choose on capability and licence; the memory budget is identical.

DeepSeek-V3 671B (37B active)

More in the DeepSeek-V3 family

DeepSeek-V3 671B (37B active)