Know What
Your AI Needs.
Calculate VRAM, check GPU fit, compare hardware and plan AI workloads before you buy or rent.
Three numbers
decide the fit.
The same engine runs on every page — no vendor spin, no crowd-sourced guesses.
Weights params × bits ÷ 8
The model itself. A 70B model at Q4_K_M is about 40 GB before anything else touches memory.
KV cache grows with context
Every token you feed the model is cached. At 128K context, the cache alone can rival the weights.
Overhead CUDA + fragmentation
A fixed floor of roughly 0.75 GB plus a small fraction of the weights, reserved before your first token.
Parity-tested against the JavaScript engine to the gigabyte. See the full methodology
See it at scale
Memory bends with context.
Llama 3.3 70B at Q4_K_M. Total VRAM climbs as the context window grows — the number a spec sheet never shows you.
Two ways in.
The Workspace
Type a model, pick a budget, and get the fit, the decode speed, and the buy-vs-rent call ranked across every card — on one screen.
Open the WorkspaceThe Engineering Center
Every formula, constant, and source behind the numbers — the methodology laid bare, with the reasoning you can check line by line.
Read the engineeringQuestions, answered plainly.
How much VRAM do I actually need?
Is a card with more memory always better?
My model is bigger than my card. Can I still run it?
Should I buy a GPU or rent one online?
Why are your numbers higher than the model's own page?
Stop guessing. Size it right.
Free, private, and cited to the last gigabyte.