Skip to content
Bitpute

Model guides

GPU requirements, model by model

By Bitpute · Updated 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error

One page per model, one answer per page: the VRAM it needs at every quantization, the cheapest card that fits, whether an RTX 4090 handles it, and what maximum context costs. Every number is computed from the model's published architecture — the same engine as the calculator, written up so you can read it. Coming from the hardware side instead? Browse the GPU database — the same numbers, organized by card.

Llama 3.2 1B

1.7 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Llama 3.2 3B

3.5 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Llama 3.1 8B

6.5 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Mistral 7B

6.0 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Qwen2.5 7B

5.7 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Qwen2.5 14B

11.0 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Gemma 2 9B

8.9 GB at Q4_K_M · minimum GPU: RTX 3060 12GB

Read the guide →

Gemma 2 27B

19.8 GB at Q4_K_M · minimum GPU: RTX 3090 24GB

Read the guide →

Qwen2.5 32B

22.2 GB at Q4_K_M · minimum GPU: RTX 3090 24GB

Read the guide →

DeepSeek-R1 32B

22.2 GB at Q4_K_M · minimum GPU: RTX 3090 24GB

Read the guide →

Llama 3.3 70B

45.1 GB at Q4_K_M · minimum GPU: RTX A6000 48GB

Read the guide →

Qwen2.5 72B

46.3 GB at Q4_K_M · minimum GPU: A100 80GB

Read the guide →

Missing a model you care about? Tell us — models with a public config file get added fastest.