Best local LLMs for 24 GB VRAM
Every catalogued model that fits entirely in 24 GB, with no layers moved to system RAM. Speeds are estimated on the NVIDIA GeForce RTX 4090, the most common 24 GB card in Steam's hardware survey.
- Largest popular model that fitsQwen3.6-35B-A3B36B params · Q4_K_M · needs 23.1 GB~370 tok/s Faster than you readOpen in the calculator →
- Most downloaded that fitsQwen3.8-27B28B params · Q4_K_M · needs 18.6 GB~46 tok/s Faster than you readOpen in the calculator →
“Best” is not a quality ranking. The picks are the largest widely downloaded model that fits, the largest that still answers at 30 tokens a second or more, and the most downloaded; the list puts current, widely downloaded releases first. Speeds are estimates from memory bandwidth, not benchmarks.
Is 24 GB of VRAM enough for a local LLM?
For models under 36B parameters, yes: all 211 in the catalogue fit entirely at Q4_K_M, the format most people download. None of the 51 models in the 36B and larger band fits. All 214 that fit answer at 30 tokens a second or more on the NVIDIA GeForce RTX 4090, which is faster than most people read. A longer conversation needs more memory than the 8,192 tokens counted here, and a smaller format than Q4_K_M needs less; the calculator sizes either.
| Model size | Fit in 24 GB | For example |
|---|---|---|
| Under 4B parameters | 66 of 66 | Qwen3.5-2B · needs 2.32 GB |
| 4B to 9B parameters | 58 of 58 | Qwen3.5-4B · needs 3.98 GB |
| 9B to 16B parameters | 21 of 21 | Qwen3.5-9B · needs 7.05 GB |
| 16B to 36B parameters | 66 of 66 | Qwen3.8-27B · needs 18.6 GB |
| 36B and larger parameters | 0 of 51 | none fits entirely |
Models that fit in 24 GB
| Model | Parameters | Needs | Spare | Decode | Answer |
|---|---|---|---|---|---|
| Qwen3.8-27B Qwen · Q4_K_M | 28B | 18.6 GB | 5.45 GB | ~46 tok/s | Details |
| gemma-4-26B-A4B-it Google · Q4_K_M | 26B | 16.8 GB | 7.21 GB | ~257 tok/s | Details |
| gemma-4-31B-it Google · Q4_K_M | 31B | 22.2 GB | 1.83 GB | ~33 tok/s | Details |
| Qwen3.5-9B Qwen · Q4_K_M | 9.7B | 7.05 GB | 17.0 GB | ~141 tok/s | Details |
| Qwen3.5-4B Qwen · Q4_K_M | 4.7B | 3.98 GB | 20.0 GB | ~254 tok/s | Details |
| Qwen3.6-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 0.90 GB | ~370 tok/s | Details |
| gemma-4-12B-it Google · Q4_K_M | 12B | 9.54 GB | 14.5 GB | ~84 tok/s | Details |
| Qwen3.6-27B Qwen · Q4_K_M | 28B | 18.6 GB | 5.45 GB | ~46 tok/s | Details |
| Qwen3.5-2B Qwen · Q4_K_M | 2.3B | 2.32 GB | 21.7 GB | ~556 tok/s | Details |
| gemma-4-E4B-it Google · Q4_K_M | 8.0B | 5.86 GB | 18.1 GB | ~227 tok/s | Details |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 3.71 GB | ~35 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-4B-BF16 NVIDIA · Q4_K_M | 4.0B | 3.46 GB | 20.5 GB | ~257 tok/s | Details |
| gemma-4-E2B-it Google · Q4_K_M | 5.1B | 4.02 GB | 20.0 GB | ~599 tok/s | Details |
| Qwen3.5-0.8B Qwen · Q4_K_M | 873M | 1.46 GB | 22.5 GB | ~1217 tok/s | Details |
| Qwen3-VL-8B-Instruct Qwen · Q4_K_M | 8.8B | 7.04 GB | 17.0 GB | ~116 tok/s | Details |
| Qwen3.5-27B Qwen · Q4_K_M | 28B | 18.6 GB | 5.45 GB | ~46 tok/s | Details |
| North-Micro-Vision-Instruct CohereLabs · Q4_K_M | 2.5B | 2.99 GB | 21.0 GB | ~365 tok/s | Details |
| Qwen3.5-35B-A3B Qwen · Q4_K_M | 36B | 23.1 GB | 0.90 GB | ~370 tok/s | Details |
| granite-4.2-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 16.5 GB | ~106 tok/s | Details |
| GLM-4.7-Flash zai-org · Q4_K_M | 31B | 20.4 GB | 3.59 GB | ~35 tok/s | Details |
| granite-4.1-3b IBM · Q4_K_M | 3.4B | 3.72 GB | 20.3 GB | ~247 tok/s | Details |
| LFM2.5-2.6B LiquidAI · Q4_K_M | 2.7B | 2.61 GB | 21.4 GB | ~397 tok/s | Details |
| Qwen3-VL-4B-Instruct Qwen · Q4_K_M | 4.4B | 4.72 GB | 19.3 GB | ~185 tok/s | Details |
| granite-4.1-30b IBM · Q4_K_M | 29B | 20.4 GB | 3.56 GB | ~35 tok/s | Details |
| Qwen3-VL-2B-Instruct Qwen · Q4_K_M | 2.1B | 3.05 GB | 21.0 GB | ~335 tok/s | Details |
| granite-4.2-3b IBM · Q4_K_M | 3.7B | 3.72 GB | 20.3 GB | ~247 tok/s | Details |
| LFM2.5-230M LiquidAI · Q4_K_M | 230M | 1.05 GB | 22.9 GB | ~2766 tok/s | Details |
| Qwen3-0.6B Qwen · Q4_K_M | 752M | 2.22 GB | 21.8 GB | ~515 tok/s | Details |
| gpt-oss-20b OpenAI · Q4_K_M | 21B | 13.8 GB | 10.2 GB | ~274 tok/s | Details |
| granite-4.2-30b IBM · Q4_K_M | 29B | 20.7 GB | 3.33 GB | ~35 tok/s | Details |
| granite-4.1-8b IBM · Q4_K_M | 8.8B | 7.49 GB | 16.5 GB | ~106 tok/s | Details |
| NVIDIA-Nemotron-3-Nano-30B-A3B-BF16 NVIDIA · Q4_K_M | 32B | 20.3 GB | 3.71 GB | ~35 tok/s | Details |
| LFM2.5-VL-3B LiquidAI · Q4_K_M | 3.1B | 2.85 GB | 21.1 GB | ~397 tok/s | Details |
| MiMo-V2.6-Distill-Qwen-9B XiaomiMiMo · Q4_K_M | 9.4B | 6.90 GB | 17.1 GB | ~141 tok/s | Details |
| Qwen3-4B-Instruct-2507 Qwen · Q4_K_M | 4.0B | 4.72 GB | 19.3 GB | ~185 tok/s | Details |
| Ling-3.0-tiny inclusionAI · Q4_K_M | 7.9B | 5.85 GB | 18.1 GB | ~799 tok/s | Details |
| Qwen3-8B Qwen · Q4_K_M | 8.2B | 7.04 GB | 17.0 GB | ~116 tok/s | Details |
| Olmo-3-7B-Instruct allenai · Q4_K_M | 7.3B | 7.96 GB | 16.0 GB | ~99 tok/s | Details |
| Qwen3-4B Qwen · Q4_K_M | 4.0B | 4.51 GB | 19.5 GB | ~185 tok/s | Details |
| LFM2.5-8B-A1B LiquidAI · Q4_K_M | 8.5B | 6.06 GB | 17.9 GB | ~618 tok/s | Details |
9 devices with 24 GB
The same models fit on every one of them. What changes is speed, which follows memory bandwidth: a card with twice the bandwidth decodes roughly twice as fast.
| Device | Bandwidth | Memory | Where to find one |
|---|---|---|---|
| NVIDIA GeForce RTX 4090 speeds on this page | 1008 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| AMD Radeon™ RX 7900 XTX | 960 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 3090 | 936 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| Intel Arc Pro B60 24GB | 456 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 3090 Ti | 1008 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA GeForce RTX 5090 Laptop GPU | 896 GB/s | 24 GB, dedicated | Amazon ↗ · eBay (new and used) ↗ |
| NVIDIA L4 | 300 GB/s | 24 GB, dedicated | — |
| Apple M2 | 100 GB/s | 24 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
| Apple M3 | 100 GB/s | 24 GB, unified | Amazon ↗ · eBay (new and used) ↗ |
Which card to buy, at every memory size →
Store links may pay us a commission. They never decide which models are listed — the memory arithmetic does.
What does not fit in 24 GB
Widely downloaded models that need more than 24 GB at Q4_K_M, and the smallest memory size that holds each one entirely. Each link shows what 24 GB can still do with it: a smaller format, or part of the model in system RAM at a lower speed.
| Model | Needs | Fits from | On 24 GB |
|---|---|---|---|
| Kimi-Linear-48B-A3B-Instruct | 30.4 GB | 32 GB | On 24 GB |
| Qwen3-Coder-Next | 50.0 GB | 64 GB | On 24 GB |
| NVIDIA-Nemotron-3-Super-120B-A12B-BF16 | 76.9 GB | 96 GB | On 24 GB |
| Qwen3.5-122B-A10B | 77.9 GB | 96 GB | On 24 GB |
| gpt-oss-120b | 71.9 GB | 96 GB | On 24 GB |
| Ling-3.0-flash | 78.2 GB | 96 GB | On 24 GB |
3 more models fit in 32 GB. Best local LLMs for 32 GB →