LLMBOTTLENECK.COM
217 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 32 GB VRAM

Every catalogued model that fits entirely in 32 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 5090, the most common 32 GB card in Steam's hardware survey.

Quick answer
217 catalogued models fit in 32 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is Kimi-Linear-48B-A3B-Instruct, needing 30.4 GB and running at ~652 tokens per second on the NVIDIA GeForce RTX 5090.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 32 GB of VRAM enough for a local LLM?

For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 36B and larger band, 3 of 51 fit. All 217 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 5090, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 32 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters21 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters66 of 66Qwen3.8-27B · needs 18.6 GB
36B and larger parameters3 of 51Kimi-Linear-48B-A3B-Instruct · needs 30.4 GB

Models that fit in 32 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.8-27B
Qwen · Q4_K_M
28B18.6 GB13.4 GB~82 tok/sDetails
gemma-4-26B-A4B-it
Google · Q4_K_M
26B16.8 GB15.2 GB~458 tok/sDetails
gemma-4-31B-it
Google · Q4_K_M
31B22.2 GB9.83 GB~59 tok/sDetails
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB25.0 GB~251 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB28.0 GB~451 tok/sDetails
Qwen3.6-35B-A3B
Qwen · Q4_K_M
36B23.1 GB8.90 GB~657 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB22.5 GB~149 tok/sDetails
Qwen3.6-27B
Qwen · Q4_K_M
28B18.6 GB13.4 GB~82 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB29.7 GB~988 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB26.1 GB~403 tok/sDetails
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
NVIDIA · Q4_K_M
32B20.3 GB11.7 GB~62 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB28.5 GB~458 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB28.0 GB~1064 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB30.5 GB~2164 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB25.0 GB~207 tok/sDetails
Qwen3.5-27B
Qwen · Q4_K_M
28B18.6 GB13.4 GB~82 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB29.0 GB~650 tok/sDetails
Qwen3.5-35B-A3B
Qwen · Q4_K_M
36B23.1 GB8.90 GB~657 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB24.5 GB~189 tok/sDetails
GLM-4.7-Flash
zai-org · Q4_K_M
31B20.4 GB11.6 GB~62 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB28.3 GB~440 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB29.4 GB~706 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB27.3 GB~329 tok/sDetails
granite-4.1-30b
IBM · Q4_K_M
29B20.4 GB11.6 GB~62 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB29.0 GB~596 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB28.3 GB~440 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB30.9 GB~4918 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB29.8 GB~915 tok/sDetails
gpt-oss-20b
OpenAI · Q4_K_M
21B13.8 GB18.2 GB~487 tok/sDetails
granite-4.2-30b
IBM · Q4_K_M
29B20.7 GB11.3 GB~62 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB24.5 GB~189 tok/sDetails
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
NVIDIA · Q4_K_M
32B20.3 GB11.7 GB~62 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB29.1 GB~706 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB25.1 GB~251 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB27.3 GB~329 tok/sDetails
Ling-3.0-tiny
inclusionAI · Q4_K_M
7.9B5.85 GB26.1 GB~1421 tok/sDetails
Qwen3-8B
Qwen · Q4_K_M
8.2B7.04 GB25.0 GB~207 tok/sDetails
Olmo-3-7B-Instruct
allenai · Q4_K_M
7.3B7.96 GB24.0 GB~176 tok/sDetails
Qwen3-4B
Qwen · Q4_K_M
4.0B4.51 GB27.5 GB~329 tok/sDetails
LFM2.5-8B-A1B
LiquidAI · Q4_K_M
8.5B6.06 GB25.9 GB~1100 tok/sDetails
177 more fit. The NVIDIA GeForce RTX 5090 page lists every one.All 217 models

10 devices with 32 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
NVIDIA GeForce RTX 5090
speeds on this page
1792 GB/s32 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon PRO W7800576 GB/s32 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
AMD Radeon™ AI PRO R9700640 GB/s32 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc Pro B65608 GB/s32 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Intel Arc Pro B70608 GB/s32 GB, dedicatedAmazon ↗ · eBay (new and used) ↗
Apple M1 Pro200 GB/s32 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M2 Pro200 GB/s32 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M4120 GB/s32 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M5153 GB/s32 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M6170 GB/s32 GB, unifiedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 32 GB

Widely downloaded models that need more than 32 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 32 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

9 more models fit in 48 GB. Best local LLMs for 48 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135