GPU guide
RX 7900 XT for AI: what it actually runs
20 GB of VRAM at 800 GB/s — a strong value pick for running local chat and coding models, as long as you're comfortable outside NVIDIA's CUDA ecosystem.
AI suitability
At Q4_K_M with an 8K context this card comfortably fits 8 of the 12 models in our reference list; at the near-lossless Q8_0, 7. The largest comfortable Q4 fit is Gemma 2 27B. For LLM inference the number that matters most is memory bandwidth, because generating each token means reading the whole model — on the RX 7900 XT that puts the theoretical ceiling around 177 tokens/second on Llama 3.1 8B at Q4_K_M, with real-world throughput below that.
Software & ecosystem — the AMD caveat
AMD cards run local LLMs through ROCm and the Vulkan llama.cpp backend. For plain text generation — Ollama, LM Studio, llama.cpp — support on RDNA 3 is now solid, and the tokens-per-second ceiling below is real. Where you feel the gap is CUDA-only tooling: some fine-tuning stacks, a few ComfyUI custom nodes, and bleeding-edge research code assume NVIDIA. Budget extra setup time for image/video generation and training, and prefer Linux for the smoothest ROCm experience.
Which models fit the RX 7900 XT?
Computed at an 8K context (or the model's own cap). ✓ fits comfortably (≤95% of 20 GB) · ~ tight · ✗ doesn't fit. Every model links to its own guide.
| Model | Q4_K_M | Q8_0 |
|---|---|---|
| Llama 3.2 1B | 1.6 GB ✓ | 2.0 GB ✓ |
| Llama 3.2 3B | 3.4 GB ✓ | 4.7 GB ✓ |
| Mistral 7B | 5.9 GB ✓ | 9.0 GB ✓ |
| Qwen2.5 7B | 5.3 GB ✓ | 8.5 GB ✓ |
| Llama 3.1 8B | 6.5 GB ✓ | 10.1 GB ✓ |
| Gemma 2 9B | 7.4 GB ✓ | 11.4 GB ✓ |
| Qwen2.5 14B | 10.5 GB ✓ | 16.8 GB ✓ |
| Gemma 2 27B | 19.6 GB ~ | 31.7 GB ✗ |
| Qwen2.5 32B | 21.7 GB ✗ | 36.0 GB ✗ |
| DeepSeek-R1 32B | 21.7 GB ✗ | 36.0 GB ✗ |
| Llama 3.3 70B | 44.7 GB ✗ | 76.0 GB ✗ |
| Qwen2.5 72B | 45.9 GB ✗ | 78.1 GB ✗ |
Electricity
The board is rated at 315 W. Run it under sustained load 8 hours a day and that's about 77 kWh a month — roughly $9/month at an example rate of $0.12/kWh (set your own tariff in the workspace). Idle and light chat draw far less; the figure above is the worst case, not the typical bill.
Alternatives
16 GB · 288 GB/s · 190 W
Step down — cheaper 16 GB if you don't need the extra room or speed.
See this card →24 GB · 960 GB/s · 355 W
Step up — 24 GB unlocks the 32B class and 70B with offload.
See this card →24 GB · 936 GB/s · 350 W
The used-NVIDIA value pick at a similar bandwidth, plus 4 GB and CUDA.
See this card →Compare any two of these head-to-head — speed on the same model, cost, power — in GPU Compare, or get a pick for your budget in the Recommendation Wizard.
Sizing something specific? Common mistakes when sizing VRAM covers the traps — including why a card with more memory can be the slower one.
New to GPU specifications? What actually matters, in what order.
Common questions
Is the RX 7900 XT good for AI and local LLMs?
Within its tier and ecosystem: with 20 GB of VRAM and 800 GB/s of bandwidth it fits 8 of the 12 models in our list at Q4_K_M with an 8K context, the largest being Gemma 2 27B. It runs text-generation models well through ROCm and Vulkan; CUDA-only fine-tuning and image tooling need extra setup.
How much electricity does a RX 7900 XT use?
The board is rated at 315 W. Under sustained load 8 hours a day that is about 77 kWh a month — roughly $9/month at $0.12/kWh. Idle draw is far lower.
Consumer cards, ranked by VRAM
Gaming cards — the cheapest route to VRAM. No ECC memory, and on NVIDIA no NVLink from Ada onward. Every card in this tier, smallest memory first — the one you are reading is highlighted. Cards shown without a memory figure are not in our calculation engine yet.