Skip to content
Bitpute
ToolsModelsGPUsCloudLearn

The Bitpute Index

GPU scores you
won't find on a
spec sheet.

Everyone lists VRAM and TDP. We compute what they mean: how efficiently a card uses its memory, what a token of throughput actually costs, and how good it really is for running models locally. Derived, opinionated, reproducible.

Computed, not copied

Five proprietary scores from one open method.

One reference load

Every card judged on the same model.

Ranked, not listed

Sort by what you actually care about.

Rank bylive

Ranked by overall index

Scores are normalized 0–100 across the cards shown, so filtering by class re-scales the field. Throughput uses the reference model at Q4_K_M as a bandwidth-bound ceiling — a fair yardstick, not a benchmark.

The methodology

Every score, explained.

These aren't scraped numbers. Each is derived from the raw specs with a fixed formula so you can reproduce — and argue with — every result.

VRAM efficiency

bandwidth ÷ VRAM — how many times per second a card can read its entire memory. A card with huge VRAM but modest bandwidth can hold a big model but can't feed it, so it scores low. This is the number that separates a serving card from a storage locker.

Price per GB

price ÷ VRAM — the true cost of the thing you're actually buying for inference: memory. Cuts through MSRP noise; a cheap 24 GB card often beats a flagship on the metric that decides what you can load.

Tokens/sec per $

reference tok/s ÷ (price / 1,000) — throughput you get per $1,000 of card, on one model every card can run. This is where used and last-gen cards quietly win.

Performance per watt

reference tok/s ÷ (TDP / 100) — throughput per 100 W. The metric that decides your electricity bill and whether a card belongs in a home or a datacenter.

Local Inference Score — the flagship

A weighted blend tuned for running LLMs at home or on a workstation: capability (VRAM, 25%) + speed (bandwidth, 20%) + affordability (inverse price, 25%) + desktop-fit (30%). That last axis is why a passively-cooled H100 — a phenomenal card — ranks below an RTX 4090 for local use: you can't put one in your desk. It's an opinion, stated in numbers, and you can see exactly how it's built.

Found your card?

Now put it to work.

Take the top-ranked card into the tools that use these same numbers — will it run your model, and does owning beat cloud.

Evidence & method

How this calculation works

Ranks GPUs for AI work from published specifications — VRAM, memory bandwidth, TDP and price class — via the documented methodology.

Data sources

  • NVIDIA, AMD & Intel GPU documentation
  • Bitpute Methodology

Assumptions

  • Ranked from published specifications: VRAM, memory bandwidth, TDP and price class
  • Efficiency proxies are derived (e.g. bandwidth-bound decode ceiling = bandwidth ÷ weight size)
  • Single GPU, inference-oriented unless stated

Limitations

  • Derived proxies are not a substitute for measured benchmarks on your workload
  • Real performance depends on framework, model, batch, context and drivers
  • Rankings shift as prices and availability change
  • For comparison and orientation, not a performance guarantee

Data status: Hardware specs — vendor documentation. Rankings — computed, not measured.

Why does this estimate differ from other calculators?
  • Computed proxies vs measured runs
  • Which spec dimensions are weighted
  • Price and availability snapshots
  • Rounding