Skip to content
Bitpute

How much VRAM does DBRX 132B (36B active) need?

By Bitpute · Published 12 July 2026 · Updated 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error · Figures re-checked against the calculation engine on every build

DBRX 132B (36B active) needs about 79.0 GB of VRAM at Q4_K_M, or 137.9 GB at Q8_0 and 258.9 GB at FP16. No consumer card holds it at Q4_K_M. The smallest fit in our set is the H200 141GB (141 GB). Weights plus runtime overhead; KV cache for your context length is on top.

Weights-plus-runtime footprint across the six most common quantizations. KV cache is context-dependent and comes on top. At 32K it is gigabytes on its own, before the weights; size it precisely in the GPU Memory Calculator.

This is a mixture-of-experts model: all 132B parameters must sit in memory, but only about 36B are active per token — so it loads like a 132B model and runs closer to a 36B one.

This is the base checkpoint: pretrained on raw text and not tuned to follow instructions, so it continues text rather than answering prompts. That makes it the right starting point if you plan to fine-tune your own model. For chat or assistant work, use DBRX 132B (36B active) Instruct — same parameter count, so the memory figures below are identical.

Parameters132B
Active / token36B
Context32K
VendorDatabricks
LicenseOpen weights (see card)
Released2024

Can your GPU run DBRX 132B (36B active)?

Pick your card — the answer below is computed by the same engine that produces every figure on this page.

Yes — recommended
Estimated VRAM79.0 GB
GPU VRAM141 GB
Headroom54.9 GB

H200 141GB runs DBRX 132B (36B active) at Q4_K_M with 54.9 GB to spare; comfortable at this size, where most cards cannot hold the weights at all; as a mixture-of-experts model all 132B parameters must be resident even though only 36B are active per token.

Weights plus runtime overhead. KV cache is not included — it depends on context length and this model’s published architecture.

Why
Weights (Q4_K_M)74.5 GB
Runtime overhead4.5 GB
Estimated total79.0 GB
Usable VRAM (95% of 141 GB)133.9 GB

Across every GPU we track

Estimates, not guarantees: real usage moves with runtime, driver, batch size and context. Size a specific context window in the GPU Memory Calculator.

How was this number calculated? Every figure here comes from Bitpute’s documented calculation methodology — parameters, precision, quantization, runtime overhead and usable VRAM, each formula written out in full.

Memory by quantization

QuantWeights+ runtime overhead
FP16245.9 GB258.9 GB
Q8_0130.6 GB137.9 GB
Q6_K100.8 GB106.6 GB
Q5_K_M87.4 GB92.6 GB
Q4_K_M74.5 GB79.0 GB
Q4_069.2 GB73.4 GB

Overhead = 0.75 GB + 5% of weights (CUDA context, buffers). Add KV cache on top. Weights dominate at this size, but a 32K window still costs gigabytes.

Which quantization should you actually run?

DBRX 132B (36B active) is too large for a 12–16 GB card. On H200 141GB, Q4_K_M (79.0 GB) fits with plenty of headroom for long context — the realistic single-GPU entry point. Higher quality (Q6_K 106.6 GB, FP16 258.9 GB) needs a 48 GB+ card or multiple GPUs — for occasional use, renting is usually cheaper than buying. As a mixture-of-experts model it must hold all 132B in VRAM but activates only about 36B per token, so it runs much faster than its footprint suggests — memory is the limit here, not speed.

Single-GPU fit at Q4_K_M (79.0 GB + KV)

RTX 3060 12GBRTX 4060 Ti 16GBRTX 4070 Ti SUPERRTX 3090RTX 4090RTX 5090A100 40GBRTX A6000L40SA100 80GBH100 80GBH200 141GB

Size it exactly →Rent a GPU for it →

Where the memory goes

Memory budget for DBRX 132B (36B active) on a H200 141GBStacked bar. Weights 74.5 GB, runtime overhead 4.5 GB, KV cache at 8K context 1.0 GB, against 133.9 GB usable on a H200 141GB.Memory budget for DBRX 132B (36B active) on a H200 141GB133.9 GB usable of 141 GBWeights (Q4_K_M) — 74.5 GBRuntime overhead — 4.5 GBKV cache @ 8K — 1.0 GBHeadroom — 53.9 GB

Weights and overhead are exact for Q4_K_M. The KV bar is a generic 32-layer transformer at 8K; at this scale the weights dominate, but long context still adds gigabytes — size it in the calculator is for.

Common questions

How much VRAM does DBRX 132B (36B active) need?

DBRX 132B (36B active) needs about 79.0 GB of VRAM at Q4_K_M, 137.9 GB at Q8_0, or 258.9 GB at FP16. That is weights plus about 0.75 GB of runtime overhead and 5% of weight size; KV cache is additional and depends on context length.

What GPU can run DBRX 132B (36B active)?

At Q4_K_M the smallest card in our set that fits is the H200 141GB (141 GB usable at 95%). At FP16 no single GPU in our set is sufficient.

Can DBRX 132B (36B active) run on 24 GB?

Not at Q4_K_M: it needs about 79.0 GB, more than the ~22.8 GB usable on a 24 GB card. Use a larger card, multiple GPUs, or offload.

Same memory footprint as 1 other checkpoint

DBRX 132B (36B active) has an identical parameter count to this, so every figure on this page applies to it unchanged. What varies is different checkpoint roles (instruction-tuned) — not the size. Choose on capability and licence; the memory budget is identical.

DBRX 132B (36B active) Instruct

More in the DBRX family

DBRX 132B (36B active) Instruct