Skip to content
Bitpute

Model guide

What GPU do you need for Qwen2.5 72B?

By Bitpute · Published 12 July 2026 · Updated 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error

In short

Plan for about 46.3 GB at Q4_K_M — just over the line for a 48 GB card once you leave any working margin, so the honest single-card answer is 80 GB (A100 or H100), or a multi-GPU split.

Architecture at a glance

Parameters

72.7B

Layers

80

KV heads

8

Head dim

128

Max context

128K

KV @8K

2.5 GB

Those extra 2 billion parameters over Llama 3.3 70B are exactly what push it past the 48 GB class: 45.1 GB fits an A6000 by a whisker, 46.3 GB doesn't. If you own 48 GB hardware and want this model, Q4_0 (a slightly smaller 4-bit format) squeezes it back under — the calculator shows the trade. At 128K context the total reaches about 84 GB.

VRAM by quantization (8K context)

QuantTotal VRAMCheapest GPU that fits
Q4_K_M46.3 GBA100 80GB
Q5_K_M53.8 GBA100 80GB
Q6_K61.5 GBA100 80GB
Q8_078.8 GBH200 141GB
FP16145.4 GBMulti-GPU / H200+

Weights and KV cache are exact arithmetic from the model's published config; overhead (0.75 GB + 5% of weights) is a calibrated estimate. "Fits" means at most 95% of the card. Method on the Engineering Center.

Compatible GPUs at Q4_K_M

Green fits comfortably, amber is tight, faded doesn't fit — each links to that card's full page.

RTX 3060 12GB ✗RTX 4060 Ti 16GB ✗RTX 4070 Ti SUPER ✗RTX 3090 ✗RTX 4090 ✗RTX 5090 ✗RTX A6000 ~L40S ~A100 40GB ✗A100 80GB ✓H100 80GB ✓H200 141GB ✓

Can an RTX 4090 run Qwen2.5 72B?

No. Qwen2.5 72B needs about 46.3 GB at Q4_K_M, well past the 4090's 24 GB. Split it across two cards, or step up to bigger hardware.

What about maximum context?

At the full 128K window, the KV cache grows to about 40.0 GB and the total to 83.8 GB. The KV cache is the part that grows — the weights never change. To see the exact split at any context, run this model through the GPU memory calculator, check an Ollama tag in the Ollama calculator, or size a fine-tune in the training memory calculator.

Running it in the cloud

For rented hardware the sensible floor is the A100 80GB — the smallest datacenter card that holds this model comfortably at Q4_K_M. No consumer card holds it alone, so cloud rental means datacenter instances or a multi-GPU node. Hourly prices move weekly, so we don't print them here — the workspace carries the current figures and weighs rental against electricity for this exact model.

Qwen2.5 32B

32.8B params · 22.2 GB at Q4_K_M

Same family — the natural size step.

See requirements →
Llama 3.3 70B

70.6B params · 45.1 GB at Q4_K_M

Closest size in another family.

See requirements →
DeepSeek-R1 32B

32.8B params · 22.2 GB at Q4_K_M

Closest size in another family.

See requirements →
← Llama 3.3 70B · All models · Llama 3.2 1B →

Exact memory figures for every quantization: How much VRAM does Qwen 2.5 72B Instruct need?

Common questions

What is the minimum GPU for Qwen2.5 72B?

At the default Q4_K_M quantization with an 8K context, Qwen2.5 72B needs about 46.3 GB of VRAM, so the practical minimum is a A100 80GB. Weights and KV cache are exact arithmetic from the model's config; a small runtime overhead estimate is included.

How much VRAM does Qwen2.5 72B need at maximum context?

At the full 128K window, the KV cache grows to about 40.0 GB and the total to 83.8 GB.

Can an RTX 4090 run Qwen2.5 72B?

No. Qwen2.5 72B needs about 46.3 GB at Q4_K_M, well past the 4090's 24 GB. Split it across two cards, or step up to bigger hardware.

Before you buy

Compare the shortlisted cards head to head in GPU Compare, size the exact context you need in the GPU Memory Calculator, and read which quantization to run before committing to a card — dropping one format down often removes the need for the next tier up.