LLMBOTTLENECK.COM
70 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 4 GB VRAM

Every catalogued model that fits entirely in 4 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce GTX 1650, the most common 4 GB card in Steam's hardware survey.

Quick answer
70 catalogued models fit in 4 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is Qwen3.5-4B, needing 3.98 GB and running at ~32 tokens per second on the NVIDIA GeForce GTX 1650.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 4 GB of VRAM enough for a local LLM?

For models under 4B parameters, yes: 63 of the 66 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 4B to 9B band, 7 of 58 fit, and nothing larger does. 67 of the 70 that fit answer at 30 tokens a second or more on the NVIDIA GeForce GTX 1650, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 4 GBFor example
Under 4B parameters63 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters7 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters0 of 21none fits entirely
16B to 36B parameters0 of 66none fits entirely
36B and larger parameters0 of 51none fits entirely

Models that fit in 4 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB0.02 GB~32 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB1.68 GB~71 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB0.54 GB~33 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB2.54 GB~155 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB1.01 GB~46 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB0.28 GB~31 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB1.39 GB~50 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB0.95 GB~43 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB0.28 GB~31 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB2.95 GB~352 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB1.78 GB~65 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB1.15 GB~50 tok/sDetails
LFM2.5-350M
LiquidAI · Q4_K_M
354M1.13 GB2.87 GB~227 tok/sDetails
Hy-MT2-1.8B
Tencent · Q4_K_M
2.0B2.47 GB1.53 GB~52 tok/sDetails
LFM2.5-1.2B-Instruct
LiquidAI · Q4_K_M
1.2B1.63 GB2.37 GB~84 tok/sDetails
Qwen3-1.7B
Qwen · Q4_K_M
2.0B3.02 GB0.98 GB~43 tok/sDetails
Nemotron-3.5-Content-Safety
NVIDIA · Q4_K_M
4.3B3.73 GB0.27 GB~31 tok/sDetails
SmolLM3-3B
HuggingFaceTB · Q4_K_M
3.1B3.46 GB0.54 GB~35 tok/sDetails
Ministral-3-3B-Instruct-2512
Mistral AI · Q4_K_M
3.8B3.82 GB0.18 GB~29 tok/sDetails
Qwen3-1.7B-Base
Qwen · Q4_K_M
1.7B3.02 GB0.98 GB~43 tok/sDetails
Qwen2.5-VL-3B-Instruct
Qwen · Q4_K_M
3.8B3.41 GB0.59 GB~39 tok/sDetails
Qwen2.5-0.5B-Instruct
Qwen · Q4_K_M
494M1.31 GB2.69 GB~203 tok/sDetails
Qwen2.5-1.5B-Instruct
Qwen · Q4_K_M
1.5B2.15 GB1.85 GB~72 tok/sDetails
OLMo-2-0425-1B
allenai · Q4_K_M · 4,096 ctx
1.5B2.27 GB1.73 GB~64 tok/sDetails
SmolLM3-3B-Base
HuggingFaceTB · Q4_K_M
3.1B3.46 GB0.54 GB~35 tok/sDetails
DeepSeek-R1-Distill-Qwen-1.5B
DeepSeek · Q4_K_M
1.8B2.15 GB1.85 GB~72 tok/sDetails
Qwen2.5-3B-Instruct
Qwen · Q4_K_M
3.1B3.21 GB0.79 GB~39 tok/sDetails
SmolLM2-135M-Instruct
HuggingFaceTB · Q4_K_M
135M1.09 GB2.91 GB~316 tok/sDetails
Phi-tiny-MoE-instruct
Microsoft · Q4_K_M · 4,096 ctx
3.8B3.22 GB0.78 GB~102 tok/sDetails
SmolLM2-135M
HuggingFaceTB · Q4_K_M
135M1.09 GB2.91 GB~316 tok/sDetails
Qwen2.5-0.5B
Qwen · Q4_K_M
494M1.31 GB2.69 GB~203 tok/sDetails
gemma-3-4b-it
Google · Q4_K_M
4.3B3.58 GB0.42 GB~33 tok/sDetails
Qwen2.5-Coder-3B-Instruct
Qwen · Q4_K_M
3.1B3.21 GB0.79 GB~39 tok/sDetails
granite-4.0-h-micro
IBM · Q4_K_M
3.2B2.89 GB1.11 GB~42 tok/sDetails
phi-2
Microsoft · Q4_K_M · 2,048 ctx
2.8B3.18 GB0.82 GB~31 tok/sDetails
SmolLM2-360M
HuggingFaceTB · Q4_K_M
362M1.39 GB2.61 GB~155 tok/sDetails
PowerMoE-3b
IBM · Q4_K_M · 4,096 ctx
3.4B3.13 GB0.87 GB~106 tok/sDetails
gemma-3-1b-it
Google · Q4_K_M
1000M1.66 GB2.34 GB~128 tok/sDetails
Qwen2.5-Coder-1.5B-Instruct
Qwen · Q4_K_M
1.5B2.15 GB1.85 GB~72 tok/sDetails
Llama-3.2-1B-Instruct
Meta · Q4_K_M
1.2B2.02 GB1.98 GB~81 tok/sDetails
30 more fit. The NVIDIA GeForce GTX 1650 page lists every one.All 70 models

2 devices with 4 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
NVIDIA GeForce GTX 1650
speeds on this page
128.1 GB/s4 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3050 Ti Laptop GPU192 GB/s4 GB, dedicatedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 4 GB

Widely downloaded models that need more than 4 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 4 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

17 more models fit in 6 GB. Best local LLMs for 6 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135