Skip to content Model Database
506 models. One question: will it fit? Maintained by Bitpute Engineering · Updated 6 August 2026 · How we estimate · Sources · Editorial policy · Version history · Report an error
Every profile carries parameters, context window, license, MoE active weights, and VRAM computed at the site’s standard quantizations — the same math as the calculators. Reviewed July 2026.
All families Airavata Arctic Aya 23 Aya Expanse BGE Embeddings BLOOM Baichuan Baichuan 2 Cerebras-GPT ChatGLM Code Llama CodeGemma Codestral Command A Command R Command R+ DBRX DeepSeek Coder DeepSeek Coder V2 DeepSeek LLM DeepSeek Math DeepSeek-R1 DeepSeek-V2 DeepSeek-V2.5 DeepSeek-V3 Devstral Dolly v2 E5 Embeddings ERNIE 4.5 EXAONE 3.5 Falcon Falcon 2 Falcon 3 Falcon Mamba Falcon-H1 Flan-T5 Fuyu GLM-4 GLM-4.5 GPT-Neo/J/NeoX GPT-OSS GTE Embeddings Gemma Gemma 1.1 Gemma 2 Gemma 3 Gemma 3n Granite 3.0 Granite 3.1 Granite 3.3 Grok H2O Danube Hermes Hunyuan Idefics2 InternLM InternLM2 InternLM2.5 InternLM3 Jamba Janus-Pro Jina Embeddings Kimi K2 LLaVA 1.5 LLaVA-NeXT Llama 1 Llama 2 Llama 3 Llama 3.1 Llama 3.2 Llama 3.2 Vision Llama 3.3 Llama 4 MPT Magistral MedGemma MiniCPM MiniMax-Text-01 Ministral Mistral 7B Mistral Large Mistral NeMo Mistral Small Mixtral Molmo Nemotron Nomic Embed OLMo OLMo 2 OPT OpenChat OpenHathi OpenLLaMA Orca 2 PaliGemma Phi Phi-3 Phi-3.5 Phi-4 Pixtral Pythia QwQ Qwen Qwen 1.5 Qwen 2 Qwen 2 VL Qwen 2.5 Qwen 2.5 Coder Qwen 2.5 Math Qwen 2.5 VL Qwen 3 Qwen 3 Coder Qwen 3 MoE R1 Distill RecurrentGemma RedPajama-INCITE SOLAR Sarvam Seed-OSS ShieldGemma SmolLM SmolLM2 SmolLM3 Stable Code StableLM StarCoder StarCoder2 Starling TinyLlama Tülu 3 Vicuna Whisper WizardCoder WizardLM WizardMath XGen Yi Yi 1.5 Yi Coder Yi VL Zephyr mxbai Embed All tasks chat code reasoning vision math embed base audio safety Any size ≤ 4B 4–10B 10–35B 35–100B 100B+ Sort: family Params ↑ Params ↓ Newest Showing 0
Show more
Browse all 506 models Every model in the database, grouped by family. Each page shows VRAM by quantization, single-GPU fit, and a computed quantization recommendation.
Bitpute
Exact GPU memory math for local AI — every formula documented, every source cited.