Skip to content
Bitpute

GPU guide

RTX 4090 for AI: what it actually runs

By Bitpute · Published 12 July 2026 · Updated 24 July 2026 · Specs verified against manufacturer datasheets 24 July 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error

In short

24 GB VRAM 1008 GB/s bandwidth 450 W TDP consumer

The default serious-local card. 24 GB, 1008 GB/s, and enough compute that memory bandwidth — not the chip — is what limits generation speed on every model it holds.

AI suitability

It runs the whole catalogue up to DeepSeek-R1 32B and Qwen2.5 32B at Q4_K_M, with the 32B class fitting at roughly 22.2 of its 24 GB — tight but workable at an 8K context. What it cannot do, at any quantization worth using, is the 70B class alone; that takes a second card or a 48 GB step up. Plan for its 450 W under sustained load.

For LLM inference the number that matters most is memory bandwidth, because generating each token means reading the whole model. On the RTX 4090 that puts the theoretical ceiling around 207 tokens/second on Llama 3.1 8B at Q4_K_M, and about 51 tokens/second on the biggest model it comfortably holds, Qwen2.5 32B — real-world throughput lands below these ceilings. At Q4_K_M with an 8K context this card comfortably fits 10 of the 12 models in our database; at the near-lossless Q8_0, 7.

Which models fit the RTX 4090?

Computed at an 8K context (or the model's own cap). ✓ fits comfortably (≤95% of 24 GB) · ~ tight · ✗ doesn't fit. Every model links to its own guide.

ModelQ4_K_MQ8_0
Llama 3.2 1B1.7 GB 2.3 GB
Llama 3.2 3B3.5 GB 5.0 GB
Mistral 7B6.0 GB 9.3 GB
Qwen2.5 7B5.7 GB 9.1 GB
Llama 3.1 8B6.5 GB 10.1 GB
Gemma 2 9B8.9 GB 13.0 GB
Qwen2.5 14B11.0 GB 17.6 GB
Gemma 2 27B19.8 GB 31.9 GB
Qwen2.5 32B22.2 GB 36.8 GB
DeepSeek-R1 32B22.2 GB 36.8 GB
Llama 3.3 70B45.1 GB 76.6 GB
Qwen2.5 72B46.3 GB 78.8 GB

Electricity

The board is rated at 450 W. Run it under sustained load 8 hours a day and that's about 108 kWh a month — roughly $13/month at an example rate of $0.12/kWh (set your own tariff in the workspace). Idle and light chat draw far less; the figure above is the worst case, not the typical bill.

Renting instead of buying

Major clouds don't rack consumer GeForce cards, but GPU marketplaces rent the RTX 4090 directly by the hour — often the cheapest way to test before buying. If you want a datacenter instance that covers the same model set, the smallest step is the RTX A6000 (48 GB). Price the rent-vs-buy question properly in the workspace, which models cloud rental cost and electricity side by side.

Alternatives

RTX 4070 Ti SUPER

16 GB · 672 GB/s · 285 W

Step down — keeps most small models, saves money and watts.

See this card →
RTX 3090

24 GB · 936 GB/s · 350 W

Same memory tier — the fit list is identical; bandwidth and price decide.

See this card →
RTX 5090

32 GB · 1792 GB/s · 575 W

Step up — the next memory tier and what it unlocks.

See this card →

Compare any two of these head-to-head — speed on the same model, cost, power — in GPU Compare.

← RTX 3090 · All GPUs · RTX 5090 →

RTX 4090 in the Bitpute graph

Everything this card connects to — models it runs, where to rent it, when we recommend it, and its nearest rivals. Derived from the same data as the calculators; reviewed July 2026.

Runs these models

Largest fits at Q4_K_M (of 417 that fit):

Seed-OSS 36B InstructAya 23 35B ChatCommand R 35B (v01) ChatLLaVA-NeXT 34B v1.6Yi VL 34B Chat

Even at FP16 (309 fit):

Falcon 2 11BFlan-T5 XXLLlama 3.2 Vision 11B InstructBrowse all 506 model profiles →

Weights + overhead only; KV cache comes on top.

Rent it in the cloud

3 providers stock this class of card:

Vast.aiRunPodTensorDock

Typically cheapest tier: Vast.ai. On-demand this card runs about $0.44/hr.

Compare all 18 providers →

When we recommend it

Best for inference11.5–15 GB required

Ranges where our recommendation engine picks this card, given the VRAM a workload needs.

Run your own numbers →

Why bandwidth matters more than core count for token generation: memory bandwidth explained.

Capacity against speed

The two axes are independent. Position on the horizontal decides which models fit; position on the vertical decides how fast they decode once they do. Cards up and to the left of it hold less but read faster; RTX A6000 sits lower-right — more memory, less bandwidth.

VRAM against memory bandwidth, RTX 4090 highlightedScatter plot of the 12 GPUs in our set, VRAM on the horizontal axis against memory bandwidth on the vertical. RTX 4090 sits at 24 GB and 1008 GB/s.01296259238885184VRAM (GB) →Bandwidth (GB/s) →122448141RTX 4090

If a model will not load on this card

In rough order of likelihood. For reference, the largest model in our catalogue that fits this card at Q4_K_M is Seed-OSS 36B Instruct at 22.1 GB.

  1. The model is larger than usable VRAM, not nameplate VRAM. This card reports 24 GB but roughly 22.8 GB is available to a model after the display buffer, driver and CUDA context. A model sized against the nameplate figure will appear to fit and then fail.
  2. KV cache grew past the headroom. Weights are fixed; the cache is not. A load that succeeds at 2K context can fail at 32K on the same card, because the cache comes out of the same budget. If it loaded yesterday and fails today, context length is the first thing to check. See KV cache.
  3. The quantization is heavier than assumed. Q4_K_M is roughly 4.85 bits per weight, not 4. On a large model that difference is gigabytes. Confirm which file you actually downloaded — see which quantization to run.
  4. Something else is already holding VRAM. A browser with hardware acceleration, another model still resident, or a previous process that did not release memory. On Linux nvidia-smi shows what is allocated.
  5. It loaded but generation is very slow. That usually means layers were offloaded to system RAM rather than the load failing outright. The model runs, but every offloaded layer crosses PCIe on each token. Reduce context, drop a quantization level, or use a card with more VRAM.

Common questions

Is the RTX 4090 good for AI and local LLMs?

Yes, within its tier: with 24 GB of VRAM and 1008 GB/s of memory bandwidth it comfortably runs 10 of the 12 models in our database at Q4_K_M with an 8K context, the largest being Qwen2.5 32B. Memory bandwidth caps generation speed at roughly 207 tokens/second on an 8B model in theory, with real-world results lower.

What is the biggest model an RTX 4090 can run?

Qwen2.5 32B at Q4_K_M with an 8K context is the largest comfortable fit (22.2 GB of 24 GB). Larger models need a bigger card or a multi-GPU split.

How much electricity does an RTX 4090 use?

The board is rated at 450 W. Under sustained load for 8 hours a day that is about 108 kWh a month — roughly $13/month at an example rate of $0.12/kWh. Idle draw is far lower, so light interactive use costs much less.

Consumer cards, ranked by VRAM

Gaming cards — the cheapest route to VRAM. No ECC memory, and on NVIDIA no NVLink from Ada onward. Every card in this tier, smallest memory first — the one you are reading is highlighted. Cards shown without a memory figure are not in our calculation engine yet.

RTX 3060 12GB12 GBRTX 4060 Ti 16GB16 GBRTX 4070 Ti SUPER16 GBRTX 309024 GBRTX 409024 GBRTX 509032 GBIntel Arc A770 16GBIntel Arc B580RX 7600 XTRX 7900 XTRX 7900 XTX

Where it sits in the stack

The RTX 4090 is an Ada Lovelace-generation part with GDDR6X memory and 4th-generation Tensor Cores. GDDR keeps the card affordable but caps bandwidth well below the HBM used in datacentre cards — and bandwidth, not core count, is what sets decode speed once a model fits.

It has no NVLink. Multiple cards talk over PCIe 4.0 instead, routed through the CPU, so splitting one model across two of these costs noticeably more than it would on an NVLink pair. Two cards each holding their own model is the friendlier pattern here.

Its Tensor Cores handle FP16, BF16, INT8 and FP8. FP8 halves activation memory against FP16 where a runtime supports it; Ampere cards do not have it. Most local inference runs weights at 4-6 bits via GGUF quantization rather than a native Tensor Core format, so these matter most for training and for server runtimes such as TensorRT-LLM and vLLM.

Its Tensor Cores support FP8, useful where a runtime can exploit it. None of this changes the VRAM arithmetic on this page — weights and KV cache still have to fit.

On the software side it is a CUDA device like any other NVIDIA card, so PyTorch, TensorRT, vLLM, llama.cpp and Ollama all run on it unchanged. What differs between cards is not compatibility but how much fits and how fast it decodes — which is what the numbers above measure. For the quantization formats these runtimes expect, and how much room the GPU Memory Calculator says you have left, start there.

Head to head

Pre-computed comparisons against the cards it is usually weighed against:

RTX 3090 vs RTX 4090RTX 4070 Ti SUPER vs RTX 4090RTX 4090 vs H100 80GBRTX 4090 vs RTX 5090RTX 4090 vs RTX A6000