Best local LLMs for 4 GB VRAM
Every catalogued model that fits entirely in 4 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce GTX 1650, the most common 4 GB card in Steam's hardware survey.
- Largest popular model that fitsQwen3.5-4B4.7B params · Q4_K_M · needs 3.98 GB~32 tok/s Faster than you readOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 4 GB of VRAM enough for a local LLM?
For models under 4B parameters, yes: 63 of the 66 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 4B to 9B band, 7 of 58 fit, and nothing larger does. 67 of the 70 that fit answer at 30 tokens a second or more on the NVIDIA GeForce GTX 1650, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 4 GB | For example |
|---|---|---|
| Under 4B parameters | 63 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 7 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 0 of 21 | none fits entirely |
| 16B to 36B parameters | 0 of 66 | none fits entirely |
| 36B and larger parameters | 0 of 51 | none fits entirely |
Models that fit in 4 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 0.02 GB | ~32 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 1.68 GB | ~71 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 0.54 GB | ~33 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 2.54 GB | ~155 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 1.01 GB | ~46 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 0.28 GB | ~31 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 1.39 GB | ~50 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 0.95 GB | ~43 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 0.28 GB | ~31 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 2.95 GB | ~352 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 1.78 GB | ~65 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 1.15 GB | ~50 tok/s | Details |
| LFM2.5-350M LiquidAI · Q4_K_M | 354M | 1.13 GB | 2.87 GB | ~227 tok/s | Details |
| Hy-MT2-1.8B Tencent · Q4_K_M | 2.0B | 2.47 GB | 1.53 GB | ~52 tok/s | Details |
| LFM2.5-1.2B-Instruct LiquidAI · Q4_K_M | 1.2B | 1.63 GB | 2.37 GB | ~84 tok/s | Details |
| Qwen3-1.7B Qwen · Q4_K_M | 2.0B | 3.02 GB | 0.98 GB | ~43 tok/s | Details |
| Nemotron-3.5-Content-Safety NVIDIA · Q4_K_M | 4.3B | 3.73 GB | 0.27 GB | ~31 tok/s | Details |
| SmolLM3-3B HuggingFaceTB · Q4_K_M | 3.1B | 3.46 GB | 0.54 GB | ~35 tok/s | Details |
| Ministral-3-3B-Instruct-2512 Mistral AI · Q4_K_M | 3.8B | 3.82 GB | 0.18 GB | ~29 tok/s | Details |
| Qwen3-1.7B-Base Qwen · Q4_K_M | 1.7B | 3.02 GB | 0.98 GB | ~43 tok/s | Details |
| Qwen2.5-VL-3B-Instruct Qwen · Q4_K_M | 3.8B | 3.41 GB | 0.59 GB | ~39 tok/s | Details |
| Qwen2.5-0.5B-Instruct Qwen · Q4_K_M | 494M | 1.31 GB | 2.69 GB | ~203 tok/s | Details |
| Qwen2.5-1.5B-Instruct Qwen · Q4_K_M | 1.5B | 2.15 GB | 1.85 GB | ~72 tok/s | Details |
| OLMo-2-0425-1B allenai · Q4_K_M · 4,096 ctx | 1.5B | 2.27 GB | 1.73 GB | ~64 tok/s | Details |
| SmolLM3-3B-Base HuggingFaceTB · Q4_K_M | 3.1B | 3.46 GB | 0.54 GB | ~35 tok/s | Details |
| DeepSeek-R1-Distill-Qwen-1.5B DeepSeek · Q4_K_M | 1.8B | 2.15 GB | 1.85 GB | ~72 tok/s | Details |
| Qwen2.5-3B-Instruct Qwen · Q4_K_M | 3.1B | 3.21 GB | 0.79 GB | ~39 tok/s | Details |
| SmolLM2-135M-Instruct HuggingFaceTB · Q4_K_M | 135M | 1.09 GB | 2.91 GB | ~316 tok/s | Details |
| Phi-tiny-MoE-instruct Microsoft · Q4_K_M · 4,096 ctx | 3.8B | 3.22 GB | 0.78 GB | ~102 tok/s | Details |
| SmolLM2-135M HuggingFaceTB · Q4_K_M | 135M | 1.09 GB | 2.91 GB | ~316 tok/s | Details |
| Qwen2.5-0.5B Qwen · Q4_K_M | 494M | 1.31 GB | 2.69 GB | ~203 tok/s | Details |
| gemma-3-4b-it Google · Q4_K_M | 4.3B | 3.58 GB | 0.42 GB | ~33 tok/s | Details |
| Qwen2.5-Coder-3B-Instruct Qwen · Q4_K_M | 3.1B | 3.21 GB | 0.79 GB | ~39 tok/s | Details |
| granite-4.0-h-micro IBM · Q4_K_M | 3.2B | 2.89 GB | 1.11 GB | ~42 tok/s | Details |
| phi-2 Microsoft · Q4_K_M · 2,048 ctx | 2.8B | 3.18 GB | 0.82 GB | ~31 tok/s | Details |
| SmolLM2-360M HuggingFaceTB · Q4_K_M | 362M | 1.39 GB | 2.61 GB | ~155 tok/s | Details |
| PowerMoE-3b IBM · Q4_K_M · 4,096 ctx | 3.4B | 3.13 GB | 0.87 GB | ~106 tok/s | Details |
| gemma-3-1b-it Google · Q4_K_M | 1000M | 1.66 GB | 2.34 GB | ~128 tok/s | Details |
| Qwen2.5-Coder-1.5B-Instruct Qwen · Q4_K_M | 1.5B | 2.15 GB | 1.85 GB | ~72 tok/s | Details |
| Llama-3.2-1B-Instruct Meta · Q4_K_M | 1.2B | 2.02 GB | 1.98 GB | ~81 tok/s | Details |
2 devices with 4 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| NVIDIA GeForce GTX 1650 speeds on this page | 128.1 GB/s | 4 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 3050 Ti Laptop GPU | 192 GB/s | 4 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 4 GB
Widely downloaded models that need more than 4 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 4 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 4 GB |
|---|---|---|---|
| gemma-4-E4B-it | 5.86 GB | 6 GB | On 4 GB |
| gemma-4-E2B-it | 4.02 GB | 6 GB | On 4 GB |
| Qwen3-VL-4B-Instruct | 4.72 GB | 6 GB | On 4 GB |
| Qwen3-4B-Instruct-2507 | 4.72 GB | 6 GB | On 4 GB |
| Ling-3.0-tiny | 5.85 GB | 6 GB | On 4 GB |
| Qwen3-4B | 4.51 GB | 6 GB | On 4 GB |
17 more models fit in 6 GB. Best local LLMs for 6 GB →