Best local LLMs for 48 GB VRAM
Every catalogued model that fits entirely in 48 GB, with no layers moved to system RAM. Speeds are estimated on the AMD Radeon PRO W7900.
- Largest popular model that fitsKimi-Linear-48B-A3B-Instruct49B params · Q4_K_M · needs 30.4 GB~360 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.8-27B28B params · Q4_K_M · needs 18.6 GB~45 tok/s Faster than you readOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 48 GB of VRAM enough for a local LLM?
For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. In the 36B and larger band, 12 of 51 fit. 216 of the 226 that fit answer at 30 tokens a second or more on the AMD Radeon PRO W7900, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 48 GB | For example |
|---|---|---|
| Under 4B parameters | 66 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 58 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 21 of 21 | Qwen3.5-9B · needs 7.05 GB |
| 16B to 36B parameters | 66 of 66 | Qwen3.8-27B · needs 18.6 GB |
| 36B and larger parameters | 12 of 51 | Kimi-Linear-48B-A3B-Instruct · needs 30.4 GB |
Models that fit in 48 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.8-27B Qwen · Q4_K_M | 28B | 18.6 GB | 29.4 GB | ~45 tok/s | Details |
| gemma-4-26B-A4B-it Google · Q4_K_M | 26B | 16.8 GB | 31.2 GB | ~253 tok/s | Details |
| gemma-4-31B-it Google · Q4_K_M | 31B | 22.2 GB | 25.8 GB | ~33 tok/s | Details |
| Qwen3.5-9B Qwen · Q4_K_M | 9.7B | 7.05 GB | 41.0 GB | ~139 tok/s | Details |
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 44.0 GB | ~249 tok/s | Details |
| Qwen3.6-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 24.9 GB | ~363 tok/s | Details |
| gemma-4-12B-it Google · Q4_K_M | 12B | 9.54 GB | 38.5 GB | ~82 tok/s | Details |
| Qwen3.6-27B Qwen · Q4_K_M | 28B | 18.6 GB | 29.4 GB | ~45 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 45.7 GB | ~546 tok/s | Details |
| gemma-4-E4B-it Google · Q4_K_M | 8.0B | 5.86 GB | 42.1 GB | ~223 tok/s | Details |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 27.7 GB | ~35 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 44.5 GB | ~253 tok/s | Details |
| gemma-4-E2B-it Google · Q4_K_M | 5.1B | 4.02 GB | 44.0 GB | ~588 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 46.5 GB | ~1196 tok/s | Details |
| Qwen3-VL-8B-Instruct Qwen · Q4_K_M | 8.8B | 7.04 GB | 41.0 GB | ~114 tok/s | Details |
| Qwen3.5-27B Qwen · Q4_K_M | 28B | 18.6 GB | 29.4 GB | ~45 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 45.0 GB | ~359 tok/s | Details |
| Qwen3.5-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 24.9 GB | ~363 tok/s | Details |
| granite-4.2-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 40.5 GB | ~104 tok/s | Details |
| GLM-4.7-Flash zai-org · Q4_K_M | 31B | 20.4 GB | 27.6 GB | ~34 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 44.3 GB | ~243 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 45.4 GB | ~390 tok/s | Details |
| Qwen3-VL-4B-Instruct Qwen · Q4_K_M | 4.4B | 4.72 GB | 43.3 GB | ~182 tok/s | Details |
| granite-4.1-30b IBM · Q4_K_M | 29B | 20.4 GB | 27.6 GB | ~34 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 45.0 GB | ~329 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 44.3 GB | ~243 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 46.9 GB | ~2717 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 45.8 GB | ~506 tok/s | Details |
| gpt-oss-20b OpenAI · Q4_K_M | 21B | 13.8 GB | 34.2 GB | ~269 tok/s | Details |
| granite-4.2-30b IBM · Q4_K_M | 29B | 20.7 GB | 27.3 GB | ~34 tok/s | Details |
| granite-4.1-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 40.5 GB | ~104 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 27.7 GB | ~35 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 45.1 GB | ~390 tok/s | Details |
| MiMo-V2.6-Distill-Qwen-9B XiaomiMiMo · Q4_K_M | 9.4B | 6.90 GB | 41.1 GB | ~139 tok/s | Details |
| Qwen3-4B-Instruct-2507 Qwen · Q4_K_M | 4.0B | 4.72 GB | 43.3 GB | ~182 tok/s | Details |
| Ling-3.0-tiny inclusionAI · Q4_K_M | 7.9B | 5.85 GB | 42.1 GB | ~785 tok/s | Details |
| Qwen3-8B Qwen · Q4_K_M | 8.2B | 7.04 GB | 41.0 GB | ~114 tok/s | Details |
| Olmo-3-7B-Instruct allenai · Q4_K_M | 7.3B | 7.96 GB | 40.0 GB | ~97 tok/s | Details |
| Qwen3-4B Qwen · Q4_K_M | 4.0B | 4.51 GB | 43.5 GB | ~182 tok/s | Details |
| LFM2.5-8B-A1B LiquidAI · Q4_K_M | 8.5B | 6.06 GB | 41.9 GB | ~607 tok/s | Details |
5 devices with 48 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| AMD Radeon PRO W7900 speeds on this page | 864 GB/s | 48 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA L40S | 864 GB/s | 48 GB, dedicated | — |
| NVIDIA RTX 6000 Ada Generation | 960 GB/s | 48 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA RTX A6000 | 768 GB/s | 48 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA RTX PRO 5000 Blackwell 48GB | 1344 GB/s | 48 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 48 GB
Widely downloaded models that need more than 48 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 48 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 48 GB |
|---|---|---|---|
| Qwen3-Coder-Next | 50.0 GB | 64 GB | On 48 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 76.9 GB | 96 GB | On 48 GB |
| Qwen3.5-122B-A10B | 77.9 GB | 96 GB | On 48 GB |
| gpt-oss-120b | 71.9 GB | 96 GB | On 48 GB |
| Ling-3.0-flash | 78.2 GB | 96 GB | On 48 GB |
| Qwen3.8-Flash-Next | 111.6 GB | 128 GB | On 48 GB |
7 more models fit in 64 GB. Best local LLMs for 64 GB →