How much work can this GPU actually do?
Choose a workload, GPU, model and settings. Every figure below is either a spec-derived calculation or a bandwidth-bounded estimate — each one labelled with the assumption behind it. Measured benchmark results are shown in a separate section, and only when real, sourced data exists.
Calculated & bounded estimates
Real benchmark results
for the exact configuration aboveHow this result was produced
VRAM requirement uses model weights, KV cache and engine overhead — derived directly from verified model and hardware specifications.
The decode ceiling is GPU memory bandwidth ÷ active model weights: a physical upper bound on single-stream decoding, not a measured rate. Concurrency is a VRAM capacity bound, not a throughput figure.
Actual throughput depends on runtime, kernels, batching, context, workload and system configuration. Measured numbers appear in the section above only when sourced benchmark data exists for the exact configuration.