LLMBOTTLENECK.COM
265 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 192 GB VRAM

Every catalogued model that fits entirely in 192 GB, with no layers moved to system RAM. Speeds are estimated on the AMD Instinct MI300X Accelerator.

Quick answer
265 catalogued models fit in 192 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is DeepSeek-V4-Flash-Vision-Exp, needing 188.1 GB and running at ~367 tokens per second on the AMD Instinct MI300X Accelerator.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 192 GB of VRAM enough for a local LLM?

Yes: every model in the catalogue fits entirely at Q4_K_M, the format most people download. All 265 that fit answer at 30 tokens a second or more on the AMD Instinct MI300X Accelerator, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 192 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters21 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters66 of 66Qwen3.8-27B · needs 18.6 GB
36B and larger parameters51 of 51DeepSeek-V4-Flash-0731 · needs 187.5 GB

Models that fit in 192 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.8-27B
Qwen · Q4_K_M
28B18.6 GB173.4 GB~279 tok/sDetails
DeepSeek-V4-Flash-0731
DeepSeek · Q4_K_M
304B187.5 GB4.48 GB~367 tok/sDetails
Qwen3.8-Flash-Next
Qwen · Q4_K_M
180B111.6 GB80.4 GB~1148 tok/sDetails
gemma-4-26B-A4B-it
Google · Q4_K_M
26B16.8 GB175.2 GB~1551 tok/sDetails
DeepSeek-V4-Flash-Vision-Exp
DeepSeek · Q4_K_M
305B188.1 GB3.87 GB~367 tok/sDetails
gemma-4-31B-it
Google · Q4_K_M
31B22.2 GB169.8 GB~201 tok/sDetails
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB185.0 GB~850 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB188.0 GB~1529 tok/sDetails
Qwen3.6-35B-A3B
Qwen · Q4_K_M
36B23.1 GB168.9 GB~2227 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB182.5 GB~505 tok/sDetails
Inkling-Small
thinkingmachines · Q4_K_M
266B165.5 GB26.5 GB~580 tok/sDetails
Qwen3.6-27B
Qwen · Q4_K_M
28B18.6 GB173.4 GB~279 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB189.7 GB~3348 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB186.1 GB~1365 tok/sDetails
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
NVIDIA · Q4_K_M
32B20.3 GB171.7 GB~212 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB188.5 GB~1551 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB188.0 GB~3606 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB190.5 GB~7334 tok/sDetails
DeepSeek-V4-Flash
DeepSeek · Q4_K_M
284B175.9 GB16.1 GB~367 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB185.0 GB~702 tok/sDetails
MiniMax-M2.7
MiniMaxAI · Q4_K_M
229B141.2 GB50.8 GB~481 tok/sDetails
Qwen3.5-27B
Qwen · Q4_K_M
28B18.6 GB173.4 GB~279 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB189.0 GB~2201 tok/sDetails
Hy3
Tencent · Q4_K_M
299B181.2 GB10.8 GB~292 tok/sDetails
Qwen3.5-35B-A3B
Qwen · Q4_K_M
36B23.1 GB168.9 GB~2227 tok/sDetails
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
NVIDIA · Q4_K_M
124B76.9 GB115.1 GB~54 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB184.5 GB~639 tok/sDetails
GLM-4.7-Flash
zai-org · Q4_K_M
31B20.4 GB171.6 GB~210 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB188.3 GB~1491 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB189.4 GB~2391 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB187.3 GB~1115 tok/sDetails
granite-4.1-30b
IBM · Q4_K_M
29B20.4 GB171.6 GB~210 tok/sDetails
Qwen3.5-122B-A10B
Qwen · Q4_K_M
125B77.9 GB114.1 GB~815 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB189.0 GB~2021 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB188.3 GB~1491 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB190.9 GB~16667 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB189.8 GB~3101 tok/sDetails
Qwen3-Coder-Next
Qwen · Q4_K_M
80B50.0 GB142.0 GB~1481 tok/sDetails
gpt-oss-20b
OpenAI · Q4_K_M
21B13.8 GB178.2 GB~1651 tok/sDetails
MiniMax-M2.5
MiniMaxAI · Q4_K_M
229B141.2 GB50.8 GB~481 tok/sDetails
225 more fit. The AMD Instinct MI300X Accelerator page lists every one.All 265 models

2 devices with 192 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
AMD Instinct MI300X Accelerator
speeds on this page
5300 GB/s192 GB, dedicated—
Apple M2 Ultra800 GB/s192 GB, unifiedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

llmbottleneck
catalogue 2026-10-03models 327devices 135