LLMBOTTLENECK.COM
130 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 8 GB VRAM

Every catalogued model that fits entirely in 8 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 4060 Laptop GPU, the most common 8 GB card in Steam's hardware survey.

Quick answer
130 catalogued models fit in 8 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is Olmo-3-7B-Instruct, needing 7.96 GB and running at ~25 tokens per second on the NVIDIA GeForce RTX 4060 Laptop GPU.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 8 GB of VRAM enough for a local LLM?

For models under 9B parameters, yes: 122 of the 124 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 9B to 16B band, 7 of 21 fit, and nothing larger does. 115 of the 130 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 4060 Laptop GPU, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 8 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters56 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters7 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters0 of 66none fits entirely
36B and larger parameters0 of 51none fits entirely

Models that fit in 8 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB0.95 GB~36 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB4.02 GB~64 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB5.68 GB~141 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB2.14 GB~58 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB4.54 GB~65 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB3.98 GB~152 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB6.54 GB~309 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB0.96 GB~30 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB5.01 GB~93 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB0.51 GB~27 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB4.28 GB~63 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB5.39 GB~101 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB3.28 GB~47 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB4.95 GB~85 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB4.28 GB~63 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB6.95 GB~703 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB5.78 GB~131 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB0.51 GB~27 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB5.15 GB~101 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB1.10 GB~36 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB3.28 GB~47 tok/sDetails
Ling-3.0-tiny
inclusionAI · Q4_K_M
7.9B5.85 GB2.15 GB~203 tok/sDetails
Qwen3-8B
Qwen · Q4_K_M
8.2B7.04 GB0.96 GB~30 tok/sDetails
Olmo-3-7B-Instruct
allenai · Q4_K_M
7.3B7.96 GB0.04 GB~25 tok/sDetails
Qwen3-4B
Qwen · Q4_K_M
4.0B4.51 GB3.49 GB~47 tok/sDetails
LFM2.5-8B-A1B
LiquidAI · Q4_K_M
8.5B6.06 GB1.94 GB~157 tok/sDetails
LFM2.5-350M
LiquidAI · Q4_K_M
354M1.13 GB6.87 GB~453 tok/sDetails
Hy-MT2-1.8B
Tencent · Q4_K_M
2.0B2.47 GB5.53 GB~104 tok/sDetails
LFM2.5-1.2B-Instruct
LiquidAI · Q4_K_M
1.2B1.63 GB6.37 GB~167 tok/sDetails
Qwen3-1.7B
Qwen · Q4_K_M
2.0B3.02 GB4.98 GB~85 tok/sDetails
Nemotron-3.5-Content-Safety
NVIDIA · Q4_K_M
4.3B3.73 GB4.27 GB~63 tok/sDetails
Hy-MT2-7B
Tencent · Q4_K_M
8.0B6.50 GB1.50 GB~31 tok/sDetails
GLM-4.6V-Flash
zai-org · Q4_K_M
10B7.45 GB0.55 GB~30 tok/sDetails
Qwen2.5-VL-7B-Instruct
Qwen · Q4_K_M
8.3B6.36 GB1.64 GB~36 tok/sDetails
SmolLM3-3B
HuggingFaceTB · Q4_K_M
3.1B3.46 GB4.54 GB~69 tok/sDetails
Ministral-3-8B-Instruct-2512
Mistral AI · Q4_K_M
8.9B7.14 GB0.86 GB~29 tok/sDetails
Ministral-3-3B-Instruct-2512
Mistral AI · Q4_K_M
3.8B3.82 GB4.18 GB~58 tok/sDetails
NVIDIA-Nemotron-Nano-9B-v2
NVIDIA · Q4_K_M
8.9B6.54 GB1.46 GB~30 tok/sDetails
Qwen3-VL-8B-Thinking
Qwen · Q4_K_M
8.8B7.04 GB0.96 GB~30 tok/sDetails
DeepSeek-R1-0528-Qwen3-8B
DeepSeek · Q4_K_M
8.2B7.04 GB0.96 GB~30 tok/sDetails
90 more fit. The NVIDIA GeForce RTX 4060 Laptop GPU page lists every one.All 130 models

31 devices with 8 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
NVIDIA GeForce RTX 4060 Laptop GPU
speeds on this page
256 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4060272 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3050 8GB224 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5060448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5060 Laptop GPU384 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3060 Ti (GDDR6)448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3070448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4070 Laptop GPU256 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3070 Ti608 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6600224 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 2060 SUPER448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5070 Laptop GPU384 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 2070 SUPER448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce GTX 1070256.3 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 7600288 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 2070448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3070 Laptop GPU448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce GTX 1080320.3 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6600 XT256 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5050320 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 2080448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 6650 XT280 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon RX 9060 XT 8GB320 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc A750512 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc A770 (8GB)512 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 2080 SUPER495.9 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3070 Ti Laptop GPU448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 3080 Laptop GPU448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 4060 Ti 8GB288 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5050 Laptop GPU384 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
NVIDIA GeForce RTX 5060 Ti 8GB448 GB/s8 GB, dedicatedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 8 GB

Widely downloaded models that need more than 8 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 8 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

7 more models fit in 10 GB. Best local LLMs for 10 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135