LLMBOTTLENECK.COM
172 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 20 GB VRAM

Every catalogued model that fits entirely in 20 GB, with no layers moved to system RAM. Speeds are estimated on the AMD Radeon RX 7900 XT, the most common 20 GB card in Steam's hardware survey.

Quick answer
172 catalogued models fit in 20 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is Hy-MT2-30B-A3B, needing 19.8 GB and running at ~233 tokens per second on the AMD Radeon RX 7900 XT.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 20 GB of VRAM enough for a local LLM?

For models under 16B parameters, yes: all 145 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 16B to 36B band, 25 of 66 fit, and nothing larger does. All 172 that fit answer at 30 tokens a second or more on the AMD Radeon RX 7900 XT, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 20 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters21 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters25 of 66Qwen3.8-27B · needs 18.6 GB
36B and larger parameters0 of 51none fits entirely

Models that fit in 20 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.8-27B
Qwen · Q4_K_M
28B18.6 GB1.45 GB~42 tok/sDetails
gemma-4-26B-A4B-it
Google · Q4_K_M
26B16.8 GB3.21 GB~234 tok/sDetails
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB13.0 GB~128 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB16.0 GB~231 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB10.5 GB~76 tok/sDetails
Qwen3.6-27B
Qwen · Q4_K_M
28B18.6 GB1.45 GB~42 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB17.7 GB~505 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB14.1 GB~206 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB16.5 GB~234 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB16.0 GB~544 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB18.5 GB~1107 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB13.0 GB~106 tok/sDetails
Qwen3.5-27B
Qwen · Q4_K_M
28B18.6 GB1.45 GB~42 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB17.0 GB~332 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB12.5 GB~96 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB16.3 GB~225 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB17.4 GB~361 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB15.3 GB~168 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB17.0 GB~305 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB16.3 GB~225 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB18.9 GB~2516 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB17.8 GB~468 tok/sDetails
gpt-oss-20b
OpenAI · Q4_K_M
21B13.8 GB6.24 GB~249 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB12.5 GB~96 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB17.1 GB~361 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB13.1 GB~128 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB15.3 GB~168 tok/sDetails
Ling-3.0-tiny
inclusionAI · Q4_K_M
7.9B5.85 GB14.1 GB~727 tok/sDetails
Qwen3-8B
Qwen · Q4_K_M
8.2B7.04 GB13.0 GB~106 tok/sDetails
Olmo-3-7B-Instruct
allenai · Q4_K_M
7.3B7.96 GB12.0 GB~90 tok/sDetails
Qwen3-4B
Qwen · Q4_K_M
4.0B4.51 GB15.5 GB~168 tok/sDetails
LFM2.5-8B-A1B
LiquidAI · Q4_K_M
8.5B6.06 GB13.9 GB~562 tok/sDetails
LFM2.5-350M
LiquidAI · Q4_K_M
354M1.13 GB18.9 GB~1624 tok/sDetails
Hy-MT2-1.8B
Tencent · Q4_K_M
2.0B2.47 GB17.5 GB~374 tok/sDetails
LLaDA2.0-mini
inclusionAI · Q4_K_M
16B11.0 GB9.00 GB~582 tok/sDetails
LFM2.5-1.2B-Instruct
LiquidAI · Q4_K_M
1.2B1.63 GB18.4 GB~599 tok/sDetails
Ministral-3-14B-Instruct-2512
Mistral AI · Q4_K_M
14B10.4 GB9.62 GB~68 tok/sDetails
Qwen3-1.7B
Qwen · Q4_K_M
2.0B3.02 GB17.0 GB~305 tok/sDetails
Nemotron-3.5-Content-Safety
NVIDIA · Q4_K_M
4.3B3.73 GB16.3 GB~225 tok/sDetails
Hy-MT2-7B
Tencent · Q4_K_M
8.0B6.50 GB13.5 GB~109 tok/sDetails
132 more fit. The AMD Radeon RX 7900 XT page lists every one.All 172 models

The 20 GB device in the catalogue

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
AMD Radeon RX 7900 XT
speeds on this page
800 GB/s20 GB, dedicatedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 20 GB

Widely downloaded models that need more than 20 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 20 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

42 more models fit in 24 GB. Best local LLMs for 24 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135