LLMBOTTLENECK.COM
137 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 10 GB VRAM

Every catalogued model that fits entirely in 10 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 3080, the most common 10 GB card in Steam's hardware survey.

Quick answer
137 catalogued models fit in 10 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is gemma-4-12B-it, needing 9.54 GB and running at ~63 tokens per second on the NVIDIA GeForce RTX 3080.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 10 GB of VRAM enough for a local LLM?

For models under 9B parameters, yes: all 124 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 9B to 16B band, 12 of 21 fit, and nothing larger does. All 137 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 3080, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 10 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters12 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters0 of 66none fits entirely
36B and larger parameters0 of 51none fits entirely

Models that fit in 10 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB2.95 GB~106 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB6.02 GB~191 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB0.46 GB~63 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB7.68 GB~419 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB4.14 GB~171 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB6.54 GB~194 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB5.98 GB~451 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB8.54 GB~918 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB2.96 GB~88 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB7.01 GB~276 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB2.51 GB~80 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB6.28 GB~187 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB7.39 GB~299 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB5.28 GB~140 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB6.95 GB~253 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB6.28 GB~187 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB8.95 GB~2086 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB7.78 GB~388 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB2.51 GB~80 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB7.15 GB~299 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB3.10 GB~106 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB5.28 GB~140 tok/sDetails
Ling-3.0-tiny
inclusionAI · Q4_K_M
7.9B5.85 GB4.15 GB~603 tok/sDetails
Qwen3-8B
Qwen · Q4_K_M
8.2B7.04 GB2.96 GB~88 tok/sDetails
Olmo-3-7B-Instruct
allenai · Q4_K_M
7.3B7.96 GB2.04 GB~75 tok/sDetails
Qwen3-4B
Qwen · Q4_K_M
4.0B4.51 GB5.49 GB~140 tok/sDetails
LFM2.5-8B-A1B
LiquidAI · Q4_K_M
8.5B6.06 GB3.94 GB~466 tok/sDetails
LFM2.5-350M
LiquidAI · Q4_K_M
354M1.13 GB8.87 GB~1346 tok/sDetails
Hy-MT2-1.8B
Tencent · Q4_K_M
2.0B2.47 GB7.53 GB~310 tok/sDetails
LFM2.5-1.2B-Instruct
LiquidAI · Q4_K_M
1.2B1.63 GB8.37 GB~497 tok/sDetails
Qwen3-1.7B
Qwen · Q4_K_M
2.0B3.02 GB6.98 GB~253 tok/sDetails
Nemotron-3.5-Content-Safety
NVIDIA · Q4_K_M
4.3B3.73 GB6.27 GB~186 tok/sDetails
Hy-MT2-7B
Tencent · Q4_K_M
8.0B6.50 GB3.50 GB~91 tok/sDetails
GLM-4.6V-Flash
zai-org · Q4_K_M
10B7.45 GB2.55 GB~90 tok/sDetails
Qwen2.5-VL-7B-Instruct
Qwen · Q4_K_M
8.3B6.36 GB3.64 GB~107 tok/sDetails
SmolLM3-3B
HuggingFaceTB · Q4_K_M
3.1B3.46 GB6.54 GB~206 tok/sDetails
Ministral-3-8B-Instruct-2512
Mistral AI · Q4_K_M
8.9B7.14 GB2.86 GB~86 tok/sDetails
Ministral-3-3B-Instruct-2512
Mistral AI · Q4_K_M
3.8B3.82 GB6.18 GB~171 tok/sDetails
NVIDIA-Nemotron-Nano-9B-v2
NVIDIA · Q4_K_M
8.9B6.54 GB3.46 GB~90 tok/sDetails
Qwen3-VL-8B-Thinking
Qwen · Q4_K_M
8.8B7.04 GB2.96 GB~88 tok/sDetails
97 more fit. The NVIDIA GeForce RTX 3080 page lists every one.All 137 models

2 devices with 10 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
NVIDIA GeForce RTX 3080
speeds on this page
760 GB/s10 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc B570380 GB/s10 GB, dedicatedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 10 GB

Widely downloaded models that need more than 10 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 10 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

5 more models fit in 11 GB. Best local LLMs for 11 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135