The Bitpute Index
GPU scores you
won't find on a
spec sheet.
Everyone lists VRAM and TDP. We compute what they mean: how efficiently a card uses its memory, what a token of throughput actually costs, and how good it really is for running models locally. Derived, opinionated, reproducible.
Computed, not copied
Five proprietary scores from one open method.
One reference load
Every card judged on the same model.
Ranked, not listed
Sort by what you actually care about.
—
Ranked by overall index
—Scores are normalized 0–100 across the cards shown, so filtering by class re-scales the field. Throughput uses the reference model at Q4_K_M as a bandwidth-bound ceiling — a fair yardstick, not a benchmark.
The methodology
Every score, explained.
These aren't scraped numbers. Each is derived from the raw specs with a fixed formula so you can reproduce — and argue with — every result.
VRAM efficiency
bandwidth ÷ VRAM — how many times per second a card can read its entire memory. A card with huge VRAM but modest bandwidth can hold a big model but can't feed it, so it scores low. This is the number that separates a serving card from a storage locker.
Price per GB
price ÷ VRAM — the true cost of the thing you're actually buying for inference: memory. Cuts through MSRP noise; a cheap 24 GB card often beats a flagship on the metric that decides what you can load.
Tokens/sec per $
reference tok/s ÷ (price / 1,000) — throughput you get per $1,000 of card, on one model every card can run. This is where used and last-gen cards quietly win.
Performance per watt
reference tok/s ÷ (TDP / 100) — throughput per 100 W. The metric that decides your electricity bill and whether a card belongs in a home or a datacenter.
Local Inference Score — the flagship
A weighted blend tuned for running LLMs at home or on a workstation: capability (VRAM, 25%) + speed (bandwidth, 20%) + affordability (inverse price, 25%) + desktop-fit (30%). That last axis is why a passively-cooled H100 — a phenomenal card — ranks below an RTX 4090 for local use: you can't put one in your desk. It's an opinion, stated in numbers, and you can see exactly how it's built.
Found your card?
Now put it to work.
Take the top-ranked card into the tools that use these same numbers — will it run your model, and does owning beat cloud.
Evidence & method
How this calculation works
Ranks GPUs for AI work from published specifications — VRAM, memory bandwidth, TDP and price class — via the documented methodology.
Data sources
- NVIDIA, AMD & Intel GPU documentation
- Bitpute Methodology
Assumptions
- Ranked from published specifications: VRAM, memory bandwidth, TDP and price class
- Efficiency proxies are derived (e.g. bandwidth-bound decode ceiling = bandwidth ÷ weight size)
- Single GPU, inference-oriented unless stated
Limitations
- Derived proxies are not a substitute for measured benchmarks on your workload
- Real performance depends on framework, model, batch, context and drivers
- Rankings shift as prices and availability change
- For comparison and orientation, not a performance guarantee
Data status: Hardware specs — vendor documentation. Rankings — computed, not measured.
Related
Why does this estimate differ from other calculators?
- Computed proxies vs measured runs
- Which spec dimensions are weighted
- Price and availability snapshots
- Rounding