By memory size / Q4_K_M / 8,192 tokens of context
What runs on your amount of VRAM
Whether a model fits depends only on how much memory you have. Find your number to see every model that fits and how fast it runs.
| Memory | Models that fit | Largest popular model that fits | Speeds quoted on |
|---|---|---|---|
| 4 GB VRAM | 70 | Qwen3.5-4B needs 3.98 GB | NVIDIA GeForce GTX 1650 |
| 6 GB VRAM | 87 | llama-7b needs 5.96 GB | NVIDIA GeForce RTX 3060 Laptop GPU |
| 8 GB VRAM | 130 | Olmo-3-7B-Instruct needs 7.96 GB | NVIDIA GeForce RTX 4060 Laptop GPU |
| 10 GB VRAM | 137 | gemma-4-12B-it needs 9.54 GB | NVIDIA GeForce RTX 3080 |
| 11 GB VRAM | 142 | LLaDA2.0-mini needs 11.0 GB | NVIDIA GeForce GTX 1080 Ti |
| 12 GB VRAM | 147 | Qwen2.5-Coder-14B-Instruct needs 11.4 GB | NVIDIA GeForce RTX 3060 12GB |
| 16 GB VRAM | 153 | gpt-oss-20b needs 13.8 GB | NVIDIA GeForce RTX 5060 Ti 16GB |
| 20 GB VRAM | 172 | Hy-MT2-30B-A3B needs 19.8 GB | AMD Radeon RX 7900 XT |
| 24 GB VRAM | 214 | Qwen3.6-35B-A3B needs 23.1 GB | NVIDIA GeForce RTX 4090 |
| 32 GB VRAM | 217 | Kimi-Linear-48B-A3B-Instruct needs 30.4 GB | NVIDIA GeForce RTX 5090 |
| 48 GB VRAM | 226 | Kimi-Linear-48B-A3B-Instruct needs 30.4 GB | AMD Radeon PRO W7900 |
| 64 GB unified memory | 233 | Qwen3-Coder-Next needs 50.0 GB | Apple M1 Max |
| 96 GB VRAM | 245 | Ling-3.0-flash needs 78.2 GB | NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
| 128 GB unified memory | 248 | Qwen3.8-Flash-Next needs 111.6 GB | AMD Ryzen AI Max+ 395 with Radeon 8060S |
| 192 GB VRAM | 265 | DeepSeek-V4-Flash-Vision-Exp needs 188.1 GB | AMD Instinct MI300X Accelerator |
A size appears here only when a catalogued device has exactly that much memory, so every figure on a size page is computed for a real machine. For any other amount, the calculator takes the memory you have.