Best GPU for running LLMs locally
Two numbers decide it. Memory decides which models fit at all; memory bandwidth decides how fast they answer. So start with how much memory the models you want need, then take the most bandwidth you can at that size.
The short version
| If you want | Fastest at that size | Most common at that size | Largest popular model it runs |
|---|---|---|---|
| 8 GB | RTX 3070 Ti 608 GB/s · ~70 tok/s | RTX 4060 272 GB/s · ~31 tok/s | Olmo-3-7B-Instruct needs 7.96 GB |
| 10 to 12 GB | RTX 3080 Ti 912 GB/s · ~105 tok/s | RTX 3060 12GB 360 GB/s · ~42 tok/s | Qwen2.5-Coder-14B-Instruct needs 11.4 GB |
| 16 GB | RTX 5080 960 GB/s · ~111 tok/s | RTX 5060 Ti 16GB 448 GB/s · ~52 tok/s | gpt-oss-20b needs 13.8 GB |
| 20 to 24 GB | RTX 4090 1008 GB/s · ~116 tok/s | RTX 4090 1008 GB/s · ~116 tok/s | Qwen3.6-35B-A3B needs 23.1 GB |
| 32 GB | RTX 5090 1792 GB/s · ~207 tok/s | RTX 5090 1792 GB/s · ~207 tok/s | Kimi-Linear-48B-A3B-Instruct needs 30.4 GB |
| 48 GB and more | RTX PRO 6000 Blackwell 1792 GB/s · ~207 tok/s | not in Steam’s survey | Ling-3.0-flash needs 78.2 GB |
The smallest card for the model you want
The most downloaded models of 7B and up, and the card with the least memory that holds each one whole at 8,192 tokens of context with at least 5% of the card left free. Among cards that size, the most common one is named. More memory buys a longer context or a better format; the link shows every format on that card.
8 GB
23 cards, fastest first. The RTX 3070 Ti holds 130 catalogued models entirely, up to Olmo-3-7B-Instruct among the popular ones. Every model that fits in 8 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX 3070 Ti fastest at this size | 8 GB | 608 GB/s | 130 | ~70 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc A750 | 8 GB | 512 GB/s | 130 | ~68 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc A770 (8GB) | 8 GB | 512 GB/s | 130 | ~68 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2080 SUPER | 8 GB | 495.9 GB/s | 130 | ~57 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5060 | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3060 Ti (GDDR6) | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3070 | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2060 SUPER | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2070 SUPER | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2070 | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2080 | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5060 Ti 8GB | 8 GB | 448 GB/s | 130 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| GTX 1080 | 8 GB | 320.3 GB/s | 130 | ~37 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5050 | 8 GB | 320 GB/s | 130 | ~37 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 9060 XT 8GB | 8 GB | 320 GB/s | 130 | ~42 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7600 | 8 GB | 288 GB/s | 130 | ~38 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4060 Ti 8GB | 8 GB | 288 GB/s | 130 | ~33 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6650 XT | 8 GB | 280 GB/s | 130 | ~37 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4060 most common at this size | 8 GB | 272 GB/s | 130 | ~31 tok/s | Amazon ↗ · eBay ↗ |
| GTX 1070 | 8 GB | 256.3 GB/s | 130 | ~30 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6600 XT | 8 GB | 256 GB/s | 130 | ~34 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3050 8GB | 8 GB | 224 GB/s | 130 | ~26 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6600 | 8 GB | 224 GB/s | 130 | ~30 tok/s | Amazon ↗ · eBay ↗ |
10 to 12 GB
16 cards, fastest first. The RTX 3080 Ti holds 147 catalogued models entirely, up to Qwen2.5-Coder-14B-Instruct among the popular ones. Every model that fits in 10 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX 3080 Ti fastest at this size | 12 GB | 912 GB/s | 147 | ~105 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3080 12GB | 12 GB | 912 GB/s | 147 | ~105 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3080 | 10 GB | 760 GB/s | 137 | ~88 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5070 | 12 GB | 672 GB/s | 147 | ~78 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2080 Ti | 11 GB | 616 GB/s | 142 | ~71 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4070 | 12 GB | 504 GB/s | 147 | ~58 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4070 SUPER | 12 GB | 504 GB/s | 147 | ~58 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4070 Ti | 12 GB | 504 GB/s | 147 | ~58 tok/s | Amazon ↗ · eBay ↗ |
| GTX 1080 Ti | 11 GB | 484.4 GB/s | 142 | ~56 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc B580 | 12 GB | 456 GB/s | 147 | ~60 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6750 XT | 12 GB | 432 GB/s | 147 | ~57 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7700 XT | 12 GB | 432 GB/s | 147 | ~57 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6700 XT | 12 GB | 384 GB/s | 147 | ~51 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc B570 | 10 GB | 380 GB/s | 137 | ~50 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3060 12GB most common at this size | 12 GB | 360 GB/s | 147 | ~42 tok/s | Amazon ↗ · eBay ↗ |
| RTX 2060 12GB | 12 GB | 336 GB/s | 147 | ~39 tok/s | Amazon ↗ · eBay ↗ |
16 GB
18 cards, fastest first. The RTX 5080 holds 153 catalogued models entirely, up to gpt-oss-20b among the popular ones. Every model that fits in 16 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX 5080 fastest at this size | 16 GB | 960 GB/s | 153 | ~111 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5070 Ti | 16 GB | 896 GB/s | 153 | ~104 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4080 SUPER | 16 GB | 736 GB/s | 153 | ~85 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4080 | 16 GB | 716.8 GB/s | 153 | ~83 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4070 Ti SUPER | 16 GB | 672 GB/s | 153 | ~78 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 9070 XT | 16 GB | 640 GB/s | 153 | ~85 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 9070 | 16 GB | 640 GB/s | 153 | ~85 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7800 XT | 16 GB | 624 GB/s | 153 | ~83 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6950 XT | 16 GB | 576 GB/s | 153 | ~76 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc A770 (16GB) | 16 GB | 560 GB/s | 153 | ~74 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6800 XT | 16 GB | 512 GB/s | 153 | ~68 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6800 | 16 GB | 512 GB/s | 153 | ~68 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 6900 XT | 16 GB | 512 GB/s | 153 | ~68 tok/s | Amazon ↗ · eBay ↗ |
| RTX 5060 Ti 16GB most common at this size | 16 GB | 448 GB/s | 153 | ~52 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 9060 XT 16GB | 16 GB | 320 GB/s | 153 | ~42 tok/s | Amazon ↗ · eBay ↗ |
| RTX 4060 Ti 16GB | 16 GB | 288 GB/s | 153 | ~33 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7600 XT | 16 GB | 288 GB/s | 153 | ~38 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc Pro B50 16GB | 16 GB | 224 GB/s | 153 | ~30 tok/s | Amazon ↗ · eBay ↗ |
20 to 24 GB
6 cards, fastest first. The RTX 4090 holds 214 catalogued models entirely, up to Qwen3.6-35B-A3B among the popular ones. Every model that fits in 20 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX 4090 fastest and most common at this size | 24 GB | 1008 GB/s | 214 | ~116 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3090 Ti | 24 GB | 1008 GB/s | 214 | ~116 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7900 XTX | 24 GB | 960 GB/s | 214 | ~127 tok/s | Amazon ↗ · eBay ↗ |
| RTX 3090 | 24 GB | 936 GB/s | 214 | ~108 tok/s | Amazon ↗ · eBay ↗ |
| Radeon RX 7900 XT | 20 GB | 800 GB/s | 172 | ~106 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc Pro B60 24GB | 24 GB | 456 GB/s | 214 | ~60 tok/s | Amazon ↗ · eBay ↗ |
32 GB
5 cards, fastest first. The RTX 5090 holds 217 catalogued models entirely, up to Kimi-Linear-48B-A3B-Instruct among the popular ones. Every model that fits in 32 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX 5090 fastest and most common at this size | 32 GB | 1792 GB/s | 217 | ~207 tok/s | Amazon ↗ · eBay ↗ |
| Radeon AI PRO R9700 | 32 GB | 640 GB/s | 217 | ~85 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc Pro B65 | 32 GB | 608 GB/s | 217 | ~80 tok/s | Amazon ↗ · eBay ↗ |
| Intel Arc Pro B70 | 32 GB | 608 GB/s | 217 | ~80 tok/s | Amazon ↗ · eBay ↗ |
| Radeon PRO W7800 | 32 GB | 576 GB/s | 217 | ~76 tok/s | Amazon ↗ · eBay ↗ |
48 GB and more
7 cards, fastest first. The RTX PRO 6000 Blackwell holds 245 catalogued models entirely, up to Ling-3.0-flash among the popular ones. Every model that fits in 48 GB.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell fastest at this size | 96 GB | 1792 GB/s | 245 | ~207 tok/s | Amazon ↗ · eBay ↗ |
| RTX PRO 5000 Blackwell 48GB | 48 GB | 1344 GB/s | 226 | ~155 tok/s | Amazon ↗ · eBay ↗ |
| RTX PRO 5000 Blackwell 72GB | 72 GB | 1344 GB/s | 239 | ~155 tok/s | Amazon ↗ · eBay ↗ |
| RTX 6000 Ada Generation | 48 GB | 960 GB/s | 226 | ~111 tok/s | Amazon ↗ · eBay ↗ |
| Radeon PRO W7900 | 48 GB | 864 GB/s | 226 | ~114 tok/s | Amazon ↗ · eBay ↗ |
| RTX A6000 | 48 GB | 768 GB/s | 226 | ~89 tok/s | Amazon ↗ · eBay ↗ |
| DGX Spark | 128 GB | 273 GB/s | 248 | ~32 tok/s | Amazon ↗ · eBay ↗ |
More memory than any card: Macs and mini PCs
A machine whose graphics share one large pool of memory holds models no single consumer card can, at lower bandwidth, so slower per token. Worth it when the model you want needs more than 32 GB. The memory shown is the largest configuration each is sold in.
| Card | Memory | Bandwidth | Models that fit | Qwen3-8B | Where to find one |
|---|---|---|---|---|---|
| Apple M5 Ultra | 512 GB | 1200 GB/s | 304 | ~92 tok/s | Amazon ↗ · eBay ↗ |
| Apple M3 Ultra | 512 GB | 800 GB/s | 304 | ~74 tok/s | Amazon ↗ · eBay ↗ |
| Apple M2 Ultra | 192 GB | 800 GB/s | 265 | ~74 tok/s | Amazon ↗ · eBay ↗ |
| Ryzen AI Max+ PRO 495 | 160 GB | 273.06 GB/s | 258 | ~36 tok/s | Amazon ↗ · eBay ↗ |
| Apple M1 Ultra | 128 GB | 800 GB/s | 248 | ~74 tok/s | Amazon ↗ · eBay ↗ |
| Apple M5 Max | 128 GB | 614 GB/s | 248 | ~63 tok/s | Amazon ↗ · eBay ↗ |
| Apple M4 Max | 128 GB | 546 GB/s | 248 | ~58 tok/s | Amazon ↗ · eBay ↗ |
| Apple M3 Max | 128 GB | 400 GB/s | 248 | ~47 tok/s | Amazon ↗ · eBay ↗ |
How this guide decides, and what it leaves out
- “Models that fit” counts every catalogued model that fits entirely in the card’s own memory at the standard format and 8,192 tokens of context, with no layers in system RAM.
- Speeds are estimates from memory bandwidth, calibrated against published measurements, on Qwen3-8B at Q4_K_M so every row compares the same work. The accuracy page publishes how far they are off.
- No prices. A store search has many listings and no single price this site can vouch for; the store links open that search. They may pay us a commission, and they never decide which card is named here.
- Not buying yet? Rent a card that runs your model by the hour, or compare a year of owning against renting.