Best local LLMs for 128 GB unified memory
Every catalogued model that fits entirely in 128 GB, with no layers moved to system RAM. Speeds are estimated on the AMD Ryzen AI Max+ 395 with Radeon 8060S.
- Largest popular model that fitsQwen3.8-Flash-Next180B params · Q4_K_M · needs 111.6 GB~55 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.8-27B28B params · Q4_K_M · needs 18.6 GB~13 tok/s About reading paceOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 128 GB of unified memory enough for a local LLM?
For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 36B and larger band, 34 of 51 fit. 172 of the 248 that fit answer at 30 tokens a second or more on the AMD Ryzen AI Max+ 395 with Radeon 8060S, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 128 GB | For example |
|---|---|---|
| Under 4B parameters | 66 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 58 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 21 of 21 | Qwen3.5-9B · needs 7.05 GB |
| 16B to 36B parameters | 66 of 66 | Qwen3.8-27B · needs 18.6 GB |
| 36B and larger parameters | 34 of 51 | Qwen3.8-Flash-Next · needs 111.6 GB |
Models that fit in 128 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.8-27B Qwen · Q4_K_M | 28B | 18.6 GB | 109.4 GB | ~13 tok/s | Details |
| Qwen3.8-Flash-Next Qwen · Q4_K_M | 180B | 111.6 GB | 16.4 GB | ~55 tok/s | Details |
| gemma-4-26B-A4B-it Google · Q4_K_M | 26B | 16.8 GB | 111.2 GB | ~75 tok/s | Details |
| gemma-4-31B-it Google · Q4_K_M | 31B | 22.2 GB | 105.8 GB | ~9.7 tok/s | Details |
| Qwen3.5-9B Qwen · Q4_K_M | 9.7B | 7.05 GB | 121.0 GB | ~41 tok/s | Details |
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 124.0 GB | ~74 tok/s | Details |
| Qwen3.6-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 104.9 GB | ~108 tok/s | Details |
| gemma-4-12B-it Google · Q4_K_M | 12B | 9.54 GB | 118.5 GB | ~24 tok/s | Details |
| Qwen3.6-27B Qwen · Q4_K_M | 28B | 18.6 GB | 109.4 GB | ~13 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 125.7 GB | ~162 tok/s | Details |
| gemma-4-E4B-it Google · Q4_K_M | 8.0B | 5.86 GB | 122.1 GB | ~66 tok/s | Details |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 107.7 GB | ~10 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 124.5 GB | ~75 tok/s | Details |
| gemma-4-E2B-it Google · Q4_K_M | 5.1B | 4.02 GB | 124.0 GB | ~174 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 126.5 GB | ~354 tok/s | Details |
| Qwen3-VL-8B-Instruct Qwen · Q4_K_M | 8.8B | 7.04 GB | 121.0 GB | ~34 tok/s | Details |
| Qwen3.5-27B Qwen · Q4_K_M | 28B | 18.6 GB | 109.4 GB | ~13 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 125.0 GB | ~106 tok/s | Details |
| Qwen3.5-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 104.9 GB | ~108 tok/s | Details |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 NVIDIA · Q4_K_M | 124B | 76.9 GB | 51.1 GB | ~2.6 tok/s | Details |
| granite-4.2-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 120.5 GB | ~31 tok/s | Details |
| GLM-4.7-Flash zai-org · Q4_K_M | 31B | 20.4 GB | 107.6 GB | ~10 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 124.3 GB | ~72 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 125.4 GB | ~115 tok/s | Details |
| Qwen3-VL-4B-Instruct Qwen · Q4_K_M | 4.4B | 4.72 GB | 123.3 GB | ~54 tok/s | Details |
| granite-4.1-30b IBM · Q4_K_M | 29B | 20.4 GB | 107.6 GB | ~10 tok/s | Details |
| Qwen3.5-122B-A10B Qwen · Q4_K_M | 125B | 77.9 GB | 50.1 GB | ~39 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 125.0 GB | ~98 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 124.3 GB | ~72 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 126.9 GB | ~805 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 125.8 GB | ~150 tok/s | Details |
| Qwen3-Coder-Next Qwen · Q4_K_M | 80B | 50.0 GB | 78.0 GB | ~72 tok/s | Details |
| gpt-oss-20b OpenAI · Q4_K_M | 21B | 13.8 GB | 114.2 GB | ~80 tok/s | Details |
| granite-4.2-30b IBM · Q4_K_M | 29B | 20.7 GB | 107.3 GB | ~10 tok/s | Details |
| granite-4.1-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 120.5 GB | ~31 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 107.7 GB | ~10 tok/s | Details |
| gpt-oss-120b OpenAI · Q4_K_M | 117B | 71.9 GB | 56.1 GB | ~56 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 125.1 GB | ~115 tok/s | Details |
| MiMo-V2.6-Distill-Qwen-9B XiaomiMiMo · Q4_K_M | 9.4B | 6.90 GB | 121.1 GB | ~41 tok/s | Details |
| Qwen3-4B-Instruct-2507 Qwen · Q4_K_M | 4.0B | 4.72 GB | 123.3 GB | ~54 tok/s | Details |
6 devices with 128 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| AMD Ryzen AI Max+ 395 with Radeon 8060S speeds on this page | 256 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| Apple M1 Ultra | 800 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| Apple M3 Max | 400 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| Apple M4 Max | 546 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| Apple M5 Max | 614 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA DGX Spark | 273 GB/s | 128 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 128 GB
Widely downloaded models that need more than 128 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 128 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 128 GB |
|---|---|---|---|
| DeepSeek-V4-Flash-0731 | 187.5 GB | 192 GB | On 128 GB |
| DeepSeek-V4-Flash-Vision-Exp | 188.1 GB | 192 GB | On 128 GB |
| Inkling-Small | 165.5 GB | 192 GB | On 128 GB |
| DeepSeek-V4-Flash | 175.9 GB | 192 GB | On 128 GB |
| MiniMax-M2.7 | 141.2 GB | 192 GB | On 128 GB |
| Hy3 | 181.2 GB | 192 GB | On 128 GB |
17 more models fit in 192 GB. Best local LLMs for 192 GB →