Model guides
GPU requirements, model by model
One page per model, one answer per page: the VRAM it needs at every quantization, the cheapest card that fits, whether an RTX 4090 handles it, and what maximum context costs. Every number is computed from the model's published architecture — the same engine as the calculator, written up so you can read it. Coming from the hardware side instead? Browse the GPU database — the same numbers, organized by card.
Llama 3.2 1B
1.7 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Llama 3.2 3B
3.5 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Llama 3.1 8B
6.5 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Mistral 7B
6.0 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Qwen2.5 7B
5.7 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Qwen2.5 14B
11.0 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Gemma 2 9B
8.9 GB at Q4_K_M · minimum GPU: RTX 3060 12GB
Read the guide →Gemma 2 27B
19.8 GB at Q4_K_M · minimum GPU: RTX 3090 24GB
Read the guide →Qwen2.5 32B
22.2 GB at Q4_K_M · minimum GPU: RTX 3090 24GB
Read the guide →DeepSeek-R1 32B
22.2 GB at Q4_K_M · minimum GPU: RTX 3090 24GB
Read the guide →Llama 3.3 70B
45.1 GB at Q4_K_M · minimum GPU: RTX A6000 48GB
Read the guide →Qwen2.5 72B
46.3 GB at Q4_K_M · minimum GPU: A100 80GB
Read the guide →Missing a model you care about? Tell us — models with a public config file get added fastest.