Best local LLMs for 6 GB VRAM
Every catalogued model that fits entirely in 6 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 3060 Laptop GPU, the most common 6 GB card in Steam's hardware survey.
- Largest popular model that fitsllama-7b6.7B params · Q4_K_M · needs 5.96 GB~45 tok/s Faster than you readOpen in the calculator →
- Best fast pickgemma-4-E4B-it8.0B params · Q4_K_M · needs 5.86 GB~76 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.5-4B4.7B params · Q4_K_M · needs 3.98 GB~85 tok/s Faster than you readOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 6 GB of VRAM enough for a local LLM?
For models under 4B parameters, yes: 64 of the 66 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 4B to 9B band, 23 of 58 fit, and nothing larger does. All 87 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 3060 Laptop GPU, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 6 GB | For example |
|---|---|---|
| Under 4B parameters | 64 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 23 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 0 of 21 | none fits entirely |
| 16B to 36B parameters | 0 of 66 | none fits entirely |
| 36B and larger parameters | 0 of 51 | none fits entirely |
Models that fit in 6 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 2.02 GB | ~85 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 3.68 GB | ~185 tok/s | Details |
| gemma-4-E4B-it Google · Q4_K_M | 8.0B | 5.86 GB | 0.14 GB | ~76 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 2.54 GB | ~86 tok/s | Details |
| gemma-4-E2B-it Google · Q4_K_M | 5.1B | 4.02 GB | 1.98 GB | ~200 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 4.54 GB | ~406 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 3.01 GB | ~122 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 2.28 GB | ~82 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 3.39 GB | ~132 tok/s | Details |
| Qwen3-VL-4B-Instruct Qwen · Q4_K_M | 4.4B | 4.72 GB | 1.28 GB | ~62 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 2.95 GB | ~112 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 2.28 GB | ~82 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 4.95 GB | ~922 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 3.78 GB | ~172 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 3.15 GB | ~132 tok/s | Details |
| Qwen3-4B-Instruct-2507 Qwen · Q4_K_M | 4.0B | 4.72 GB | 1.28 GB | ~62 tok/s | Details |
| Ling-3.0-tiny inclusionAI · Q4_K_M | 7.9B | 5.85 GB | 0.15 GB | ~266 tok/s | Details |
| Qwen3-4B Qwen · Q4_K_M | 4.0B | 4.51 GB | 1.49 GB | ~62 tok/s | Details |
| LFM2.5-350M LiquidAI · Q4_K_M | 354M | 1.13 GB | 4.87 GB | ~595 tok/s | Details |
| Hy-MT2-1.8B Tencent · Q4_K_M | 2.0B | 2.47 GB | 3.53 GB | ~137 tok/s | Details |
| LFM2.5-1.2B-Instruct LiquidAI · Q4_K_M | 1.2B | 1.63 GB | 4.37 GB | ~220 tok/s | Details |
| Qwen3-1.7B Qwen · Q4_K_M | 2.0B | 3.02 GB | 2.98 GB | ~112 tok/s | Details |
| Nemotron-3.5-Content-Safety NVIDIA · Q4_K_M | 4.3B | 3.73 GB | 2.27 GB | ~82 tok/s | Details |
| SmolLM3-3B HuggingFaceTB · Q4_K_M | 3.1B | 3.46 GB | 2.54 GB | ~91 tok/s | Details |
| Ministral-3-3B-Instruct-2512 Mistral AI · Q4_K_M | 3.8B | 3.82 GB | 2.18 GB | ~76 tok/s | Details |
| Qwen3-1.7B-Base Qwen · Q4_K_M | 1.7B | 3.02 GB | 2.98 GB | ~112 tok/s | Details |
| Qwen2.5-VL-3B-Instruct Qwen · Q4_K_M | 3.8B | 3.41 GB | 2.59 GB | ~103 tok/s | Details |
| Qwen3-4B-Base Qwen · Q4_K_M | 4.0B | 4.72 GB | 1.28 GB | ~62 tok/s | Details |
| Qwen2.5-0.5B-Instruct Qwen · Q4_K_M | 494M | 1.31 GB | 4.69 GB | ~534 tok/s | Details |
| Qwen2.5-7B-Instruct Qwen · Q4_K_M | 7.6B | 5.95 GB | 0.05 GB | ~47 tok/s | Details |
| Qwen2.5-1.5B-Instruct Qwen · Q4_K_M | 1.5B | 2.15 GB | 3.85 GB | ~188 tok/s | Details |
| OLMo-2-0425-1B allenai · Q4_K_M · 4,096 ctx | 1.5B | 2.27 GB | 3.73 GB | ~169 tok/s | Details |
| SmolLM3-3B-Base HuggingFaceTB · Q4_K_M | 3.1B | 3.46 GB | 2.54 GB | ~91 tok/s | Details |
| DeepSeek-R1-Distill-Qwen-1.5B DeepSeek · Q4_K_M | 1.8B | 2.15 GB | 3.85 GB | ~188 tok/s | Details |
| Qwen2.5-3B-Instruct Qwen · Q4_K_M | 3.1B | 3.21 GB | 2.79 GB | ~103 tok/s | Details |
| SmolLM2-135M-Instruct HuggingFaceTB · Q4_K_M | 135M | 1.09 GB | 4.91 GB | ~829 tok/s | Details |
| Phi-tiny-MoE-instruct Microsoft · Q4_K_M · 4,096 ctx | 3.8B | 3.22 GB | 2.78 GB | ~268 tok/s | Details |
| SmolLM2-135M HuggingFaceTB · Q4_K_M | 135M | 1.09 GB | 4.91 GB | ~829 tok/s | Details |
| Phi-4-mini-instruct Microsoft · Q4_K_M | 3.8B | 4.66 GB | 1.34 GB | ~65 tok/s | Details |
| Qwen2.5-Coder-7B-Instruct Qwen · Q4_K_M | 7.6B | 5.95 GB | 0.05 GB | ~47 tok/s | Details |
7 devices with 6 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| NVIDIA GeForce RTX 3060 Laptop GPU speeds on this page | 336 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 4050 Laptop GPU | 192 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 2060 | 336 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce GTX 1060 6GB | 192.2 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce GTX 1660 SUPER | 336 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce GTX 1660 Ti | 288 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce GTX 1660 | 192.1 GB/s | 6 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 6 GB
Widely downloaded models that need more than 6 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 6 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 6 GB |
|---|---|---|---|
| Qwen3.5-9B | 7.05 GB | 8 GB | On 6 GB |
| Qwen3-VL-8B-Instruct | 7.04 GB | 8 GB | On 6 GB |
| granite-4.2-8b | 7.49 GB | 8 GB | On 6 GB |
| granite-4.1-8b | 7.49 GB | 8 GB | On 6 GB |
| MiMo-V2.6-Distill-Qwen-9B | 6.90 GB | 8 GB | On 6 GB |
| Qwen3-8B | 7.04 GB | 8 GB | On 6 GB |
43 more models fit in 8 GB. Best local LLMs for 8 GB →