LLMBOTTLENECK.COM
153 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 16 GB VRAM

Every catalogued model that fits entirely in 16 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 5060 Ti 16GB, the most common 16 GB card in Steam's hardware survey.

Quick answer
153 catalogued models fit in 16 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is gpt-oss-20b, needing 13.8 GB and running at ~122 tokens per second on the NVIDIA GeForce RTX 5060 Ti 16GB.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 16 GB of VRAM enough for a local LLM?

For models under 16B parameters, yes: all 145 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 16B to 36B band, 6 of 66 fit, and nothing larger does. 146 of the 153 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 5060 Ti 16GB, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 16 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters21 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters6 of 66gpt-oss-20b · needs 13.8 GB
36B and larger parameters0 of 51none fits entirely

Models that fit in 16 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB8.95 GB~63 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB12.0 GB~113 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB6.46 GB~37 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB13.7 GB~247 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB10.1 GB~101 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB12.5 GB~114 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB12.0 GB~266 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB14.5 GB~541 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB8.96 GB~52 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB13.0 GB~162 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB8.51 GB~47 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB12.3 GB~110 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB13.4 GB~176 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB11.3 GB~82 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB13.0 GB~149 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB12.3 GB~110 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB14.9 GB~1230 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB13.8 GB~229 tok/sDetails
gpt-oss-20b
OpenAI · Q4_K_M
21B13.8 GB2.24 GB~122 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB8.51 GB~47 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB13.1 GB~176 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB9.10 GB~63 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB11.3 GB~82 tok/sDetails
Ling-3.0-tiny
inclusionAI · Q4_K_M
7.9B5.85 GB10.1 GB~355 tok/sDetails
Qwen3-8B
Qwen · Q4_K_M
8.2B7.04 GB8.96 GB~52 tok/sDetails
Olmo-3-7B-Instruct
allenai · Q4_K_M
7.3B7.96 GB8.04 GB~44 tok/sDetails
Qwen3-4B
Qwen · Q4_K_M
4.0B4.51 GB11.5 GB~82 tok/sDetails
LFM2.5-8B-A1B
LiquidAI · Q4_K_M
8.5B6.06 GB9.94 GB~275 tok/sDetails
LFM2.5-350M
LiquidAI · Q4_K_M
354M1.13 GB14.9 GB~794 tok/sDetails
Hy-MT2-1.8B
Tencent · Q4_K_M
2.0B2.47 GB13.5 GB~183 tok/sDetails
LLaDA2.0-mini
inclusionAI · Q4_K_M
16B11.0 GB5.00 GB~285 tok/sDetails
LFM2.5-1.2B-Instruct
LiquidAI · Q4_K_M
1.2B1.63 GB14.4 GB~293 tok/sDetails
Ministral-3-14B-Instruct-2512
Mistral AI · Q4_K_M
14B10.4 GB5.62 GB~33 tok/sDetails
Qwen3-1.7B
Qwen · Q4_K_M
2.0B3.02 GB13.0 GB~149 tok/sDetails
Nemotron-3.5-Content-Safety
NVIDIA · Q4_K_M
4.3B3.73 GB12.3 GB~110 tok/sDetails
Hy-MT2-7B
Tencent · Q4_K_M
8.0B6.50 GB9.50 GB~53 tok/sDetails
Qwen3-14B
Qwen · Q4_K_M
15B11.1 GB4.86 GB~31 tok/sDetails
GLM-4.6V-Flash
zai-org · Q4_K_M
10B7.45 GB8.55 GB~53 tok/sDetails
Qwen2.5-VL-7B-Instruct
Qwen · Q4_K_M
8.3B6.36 GB9.64 GB~63 tok/sDetails
SmolLM3-3B
HuggingFaceTB · Q4_K_M
3.1B3.46 GB12.5 GB~121 tok/sDetails
113 more fit. The NVIDIA GeForce RTX 5060 Ti 16GB page lists every one.All 153 models

21 devices with 16 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
NVIDIA GeForce RTX 5060 Ti 16GB
speeds on this page
448 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4060 Ti 16GB288 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5070 Ti896 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5080960 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon™ RX 9070 XT640 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 9060 XT 16GB320 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4070 Ti SUPER672 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4080716.8 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4080 SUPER736 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 7800 XT624 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6800 XT512 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 9070640 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6800512 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6900 XT512 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6950 XT576 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 7600 XT288 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc A770 (16GB)560 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc Pro B50 16GB224 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3080 Ti Laptop GPU512 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4090 Laptop GPU576 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5080 Laptop GPU896 GB/s16 GB, dedicatedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 16 GB

Widely downloaded models that need more than 16 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 16 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

19 more models fit in 20 GB. Best local LLMs for 20 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135