LLMBOTTLENECK.COM
248 models fit / Q4_K_M / 8,192 tokens of context

Best local LLMs for 128 GB unified memory

Every catalogued model that fits entirely in 128 GB, with no layers moved to system RAM. Speeds are estimated on the AMD Ryzen AI Max+ 395 with Radeon 8060S.

Quick answer
248 catalogued models fit in 128 GB at Q4_K_M with 8,192 tokens of context. The largest widely used one is Qwen3.8-Flash-Next, needing 111.6 GB and running at ~55 tokens per second on the AMD Ryzen AI Max+ 395 with Radeon 8060S.

“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.

Is 128 GB of unified memory enough for a local LLM?

For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 36B and larger band, 34 of 51 fit. 172 of the 248 that fit answer at 30 tokens a second or more on the AMD Ryzen AI Max+ 395 with Radeon 8060S, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.

Model sizeFit in 128 GBFor example
Under 4B parameters66 of 66Qwen3.5-2B · needs 2.32 GB
4B to 9B parameters58 of 58Qwen3.5-4B · needs 3.98 GB
9B to 16B parameters21 of 21Qwen3.5-9B · needs 7.05 GB
16B to 36B parameters66 of 66Qwen3.8-27B · needs 18.6 GB
36B and larger parameters34 of 51Qwen3.8-Flash-Next · needs 111.6 GB

Models that fit in 128 GB

ModelParametersNeedsSpareDecodeAnswer
Qwen3.8-27B
Qwen · Q4_K_M
28B18.6 GB109.4 GB~13 tok/sDetails
Qwen3.8-Flash-Next
Qwen · Q4_K_M
180B111.6 GB16.4 GB~55 tok/sDetails
gemma-4-26B-A4B-it
Google · Q4_K_M
26B16.8 GB111.2 GB~75 tok/sDetails
gemma-4-31B-it
Google · Q4_K_M
31B22.2 GB105.8 GB~9.7 tok/sDetails
Qwen3.5-9B
Qwen · Q4_K_M
9.7B7.05 GB121.0 GB~41 tok/sDetails
Qwen3.5-4B
Qwen · Q4_K_M
4.7B3.98 GB124.0 GB~74 tok/sDetails
Qwen3.6-35B-A3B
Qwen · Q4_K_M
36B23.1 GB104.9 GB~108 tok/sDetails
gemma-4-12B-it
Google · Q4_K_M
12B9.54 GB118.5 GB~24 tok/sDetails
Qwen3.6-27B
Qwen · Q4_K_M
28B18.6 GB109.4 GB~13 tok/sDetails
Qwen3.5-2B
Qwen · Q4_K_M
2.3B2.32 GB125.7 GB~162 tok/sDetails
gemma-4-E4B-it
Google · Q4_K_M
8.0B5.86 GB122.1 GB~66 tok/sDetails
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16
NVIDIA · Q4_K_M
32B20.3 GB107.7 GB~10 tok/sDetails
NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · Q4_K_M
4.0B3.46 GB124.5 GB~75 tok/sDetails
gemma-4-E2B-it
Google · Q4_K_M
5.1B4.02 GB124.0 GB~174 tok/sDetails
Qwen3.5-0.8B
Qwen · Q4_K_M
873M1.46 GB126.5 GB~354 tok/sDetails
Qwen3-VL-8B-Instruct
Qwen · Q4_K_M
8.8B7.04 GB121.0 GB~34 tok/sDetails
Qwen3.5-27B
Qwen · Q4_K_M
28B18.6 GB109.4 GB~13 tok/sDetails
North-Micro-Vision-Instruct
CohereLabs · Q4_K_M
2.5B2.99 GB125.0 GB~106 tok/sDetails
Qwen3.5-35B-A3B
Qwen · Q4_K_M
36B23.1 GB104.9 GB~108 tok/sDetails
NVIDIA-Nemotron-3-Super-120B-A12B-BF16
NVIDIA · Q4_K_M
124B76.9 GB51.1 GB~2.6 tok/sDetails
granite-4.2-8b
IBM · Q4_K_M
8.8B7.49 GB120.5 GB~31 tok/sDetails
GLM-4.7-Flash
zai-org · Q4_K_M
31B20.4 GB107.6 GB~10 tok/sDetails
granite-4.1-3b
IBM · Q4_K_M
3.4B3.72 GB124.3 GB~72 tok/sDetails
LFM2.5-2.6B
LiquidAI · Q4_K_M
2.7B2.61 GB125.4 GB~115 tok/sDetails
Qwen3-VL-4B-Instruct
Qwen · Q4_K_M
4.4B4.72 GB123.3 GB~54 tok/sDetails
granite-4.1-30b
IBM · Q4_K_M
29B20.4 GB107.6 GB~10 tok/sDetails
Qwen3.5-122B-A10B
Qwen · Q4_K_M
125B77.9 GB50.1 GB~39 tok/sDetails
Qwen3-VL-2B-Instruct
Qwen · Q4_K_M
2.1B3.05 GB125.0 GB~98 tok/sDetails
granite-4.2-3b
IBM · Q4_K_M
3.7B3.72 GB124.3 GB~72 tok/sDetails
LFM2.5-230M
LiquidAI · Q4_K_M
230M1.05 GB126.9 GB~805 tok/sDetails
Qwen3-0.6B
Qwen · Q4_K_M
752M2.22 GB125.8 GB~150 tok/sDetails
Qwen3-Coder-Next
Qwen · Q4_K_M
80B50.0 GB78.0 GB~72 tok/sDetails
gpt-oss-20b
OpenAI · Q4_K_M
21B13.8 GB114.2 GB~80 tok/sDetails
granite-4.2-30b
IBM · Q4_K_M
29B20.7 GB107.3 GB~10 tok/sDetails
granite-4.1-8b
IBM · Q4_K_M
8.8B7.49 GB120.5 GB~31 tok/sDetails
NVIDIA-Nemotron-3-Nano-30B-A3B-BF16
NVIDIA · Q4_K_M
32B20.3 GB107.7 GB~10 tok/sDetails
gpt-oss-120b
OpenAI · Q4_K_M
117B71.9 GB56.1 GB~56 tok/sDetails
LFM2.5-VL-3B
LiquidAI · Q4_K_M
3.1B2.85 GB125.1 GB~115 tok/sDetails
MiMo-V2.6-Distill-Qwen-9B
XiaomiMiMo · Q4_K_M
9.4B6.90 GB121.1 GB~41 tok/sDetails
Qwen3-4B-Instruct-2507
Qwen · Q4_K_M
4.0B4.72 GB123.3 GB~54 tok/sDetails
208 more fit. The AMD Ryzen AI Max+ 395 with Radeon 8060S page lists every one.All 248 models

6 devices with 128 GB

The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.

DeviceBandwidthMemoryWhere to find one
AMD Ryzen AI Max+ 395 with Radeon 8060S
speeds on this page
256 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M1 Ultra800 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M3 Max400 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M4 Max546 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗
Apple M5 Max614 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗
NVIDIA DGX Spark273 GB/s128 GB, unifiedAmazon ↗ · eBay (new and used) ↗

Which card to buy, at every memory size →

Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.

What does not fit in 128 GB

Widely downloaded models that need more than 128 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 128 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.

17 more models fit in 192 GB. Best local LLMs for 192 GB →

llmbottleneck
catalogue 2026-10-03models 327devices 135