Best local LLMs for 96 GB VRAM
Every catalogued model that fits entirely in 96 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA RTX PRO 6000 Blackwell Workstation Edition.
- Largest popular model that fitsLing-3.0-flash127B params · Q4_K_M · needs 78.2 GB~398 tok/s Faster than you readOpen in the calculator →
- Best fast pickMixtral-8x22B-Instruct-v0.1141B params · Q4_K_M · needs 87.7 GB~48 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.8-27B28B params · Q4_K_M · needs 18.6 GB~82 tok/s Faster than you readOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 96 GB of VRAM enough for a local LLM?
For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 36B and larger band, 31 of 51 fit. 232 of the 245 that fit answer at 30 tokens a second or more on the NVIDIA RTX PRO 6000 Blackwell Workstation Edition, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 96 GB | For example |
|---|---|---|
| Under 4B parameters | 66 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 58 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 21 of 21 | Qwen3.5-9B · needs 7.05 GB |
| 16B to 36B parameters | 66 of 66 | Qwen3.8-27B · needs 18.6 GB |
| 36B and larger parameters | 31 of 51 | NVIDIA-Nemotron-3-Super-120B-A12B-BF16 · needs 76.9 GB |
Models that fit in 96 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.8-27B Qwen · Q4_K_M | 28B | 18.6 GB | 77.4 GB | ~82 tok/s | Details |
| gemma-4-26B-A4B-it Google · Q4_K_M | 26B | 16.8 GB | 79.2 GB | ~458 tok/s | Details |
| gemma-4-31B-it Google · Q4_K_M | 31B | 22.2 GB | 73.8 GB | ~59 tok/s | Details |
| Qwen3.5-9B Qwen · Q4_K_M | 9.7B | 7.05 GB | 89.0 GB | ~251 tok/s | Details |
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 92.0 GB | ~451 tok/s | Details |
| Qwen3.6-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 72.9 GB | ~657 tok/s | Details |
| gemma-4-12B-it Google · Q4_K_M | 12B | 9.54 GB | 86.5 GB | ~149 tok/s | Details |
| Qwen3.6-27B Qwen · Q4_K_M | 28B | 18.6 GB | 77.4 GB | ~82 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 93.7 GB | ~988 tok/s | Details |
| gemma-4-E4B-it Google · Q4_K_M | 8.0B | 5.86 GB | 90.1 GB | ~403 tok/s | Details |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 75.7 GB | ~62 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 92.5 GB | ~458 tok/s | Details |
| gemma-4-E2B-it Google · Q4_K_M | 5.1B | 4.02 GB | 92.0 GB | ~1064 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 94.5 GB | ~2164 tok/s | Details |
| Qwen3-VL-8B-Instruct Qwen · Q4_K_M | 8.8B | 7.04 GB | 89.0 GB | ~207 tok/s | Details |
| Qwen3.5-27B Qwen · Q4_K_M | 28B | 18.6 GB | 77.4 GB | ~82 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 93.0 GB | ~650 tok/s | Details |
| Qwen3.5-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 72.9 GB | ~657 tok/s | Details |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 NVIDIA · Q4_K_M | 124B | 76.9 GB | 19.1 GB | ~16 tok/s | Details |
| granite-4.2-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 88.5 GB | ~189 tok/s | Details |
| GLM-4.7-Flash zai-org · Q4_K_M | 31B | 20.4 GB | 75.6 GB | ~62 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 92.3 GB | ~440 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 93.4 GB | ~706 tok/s | Details |
| Qwen3-VL-4B-Instruct Qwen · Q4_K_M | 4.4B | 4.72 GB | 91.3 GB | ~329 tok/s | Details |
| granite-4.1-30b IBM · Q4_K_M | 29B | 20.4 GB | 75.6 GB | ~62 tok/s | Details |
| Qwen3.5-122B-A10B Qwen · Q4_K_M | 125B | 77.9 GB | 18.1 GB | ~241 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 93.0 GB | ~596 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 92.3 GB | ~440 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 94.9 GB | ~4918 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 93.8 GB | ~915 tok/s | Details |
| Qwen3-Coder-Next Qwen · Q4_K_M | 80B | 50.0 GB | 46.0 GB | ~437 tok/s | Details |
| gpt-oss-20b OpenAI · Q4_K_M | 21B | 13.8 GB | 82.2 GB | ~487 tok/s | Details |
| granite-4.2-30b IBM · Q4_K_M | 29B | 20.7 GB | 75.3 GB | ~62 tok/s | Details |
| granite-4.1-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 88.5 GB | ~189 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 75.7 GB | ~62 tok/s | Details |
| gpt-oss-120b OpenAI · Q4_K_M | 117B | 71.9 GB | 24.1 GB | ~343 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 93.1 GB | ~706 tok/s | Details |
| MiMo-V2.6-Distill-Qwen-9B XiaomiMiMo · Q4_K_M | 9.4B | 6.90 GB | 89.1 GB | ~251 tok/s | Details |
| Qwen3-4B-Instruct-2507 Qwen · Q4_K_M | 4.0B | 4.72 GB | 91.3 GB | ~329 tok/s | Details |
| Ling-3.0-tiny inclusionAI · Q4_K_M | 7.9B | 5.85 GB | 90.1 GB | ~1421 tok/s | Details |
2 devices with 96 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| NVIDIA RTX PRO 6000 Blackwell Workstation Edition speeds on this page | 1792 GB/s | 96 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| Apple M2 Max | 400 GB/s | 96 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 96 GB
Widely downloaded models that need more than 96 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 96 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 96 GB |
|---|---|---|---|
| Qwen3.8-Flash-Next | 111.6 GB | 128 GB | On 96 GB |
| DeepSeek-V4-Flash-0731 | 187.5 GB | 192 GB | On 96 GB |
| DeepSeek-V4-Flash-Vision-Exp | 188.1 GB | 192 GB | On 96 GB |
| Inkling-Small | 165.5 GB | 192 GB | On 96 GB |
| DeepSeek-V4-Flash | 175.9 GB | 192 GB | On 96 GB |
| MiniMax-M2.7 | 141.2 GB | 192 GB | On 96 GB |
3 more models fit in 128 GB. Best local LLMs for 128 GB →