LLMBOTTLENECK.COM
Hugging Face downloads, last 30 days / read 2026-10-03

Most downloaded on Hugging Face

Open-weight models ordered by Hugging Face's own download count, each joined to the smallest device in this catalogue that holds it. Downloads measure popularity, not quality or use.

Where the order comes from

The ordering is Hugging Face model API’s own thirty-day download count per repository, read on 2026-10-03 and cited on each model’s page. It is not this site’s traffic, and it is not a vote — a popularity list built from our own visitors would be unfalsifiable by you and mostly noise anyway.

What it is not: Repository downloads measure attention on one host, including mirrors, CI and evaluation runs. They are not a measure of quality, and a model distributed mainly elsewhere is undercounted.

Models with a published count3270 without one
Of the top 5046run on 24 GB or less
Sizing formatQ4_K_Mor the nearest published
Reference context8,192tokens, no offload

Downloads in the last 30 days

top 50
#1 Qwen3-0.6B29.7M
#2 Qwen3-VL-8B-Instruct14.6M
#3 gemma-4-26B-A4B-it13M
#4 Qwen3-8B10.7M
#5 gemma-4-31B-it9.9M
#6 Qwen3.5-9B9M
#7 Qwen2.5-0.5B-Instruct8.7M
#8 Qwen2.5-7B-Instruct8.4M
#9 Qwen3.5-4B7.8M
#10 Qwen3-4B7.8M
#11 Qwen2.5-1.5B-Instruct7.4M
#12 Qwen3.8-27B6.9M
#13 gpt-oss-20b6.6M
#14 Qwen2.5-VL-7B-Instruct5.8M
#15 GLM-5.3-Flash5.4M
#16 Qwen3.5-2B4.9M
#17 dolphin-2.9.1-yi-1.5-34b4.7M
#18 DeepSeek-V4-Flash-07314.5M
#19 gpt-oss-120b4.5M
#20 gemma-4-E4B-it4.4M
#21 Qwen2.5-3B-Instruct4.2M
#22 Qwen3-32B4M
#23 Qwen3-4B-Instruct-25073.7M
#24 DeepSeek-V3.23.6M
#25 Qwen3-VL-4B-Instruct3.5M
#26 pythia-160m3.5M
#27 NVIDIA-Nemotron-3-Nano-4B-BF163.4M
#28 Qwen3.6-35B-A3B3.3M
#29 Qwen3-1.7B3.1M
#30 gemma-4-E2B-it3M
#31 OTel-2.0-LLM-31B-IT3M
#32 Qwen3-VL-2B-Instruct2.8M
#33 Qwen3-14B2.7M
#34 Qwen3.5-0.8B2.6M
#35 JiRackUltra_1b2.5M
#36 Qwen3.6-27B2.5M
#37 Qwen2.5-VL-3B-Instruct2.4M
#38 Qwen2.5-Coder-7B-Instruct2.3M
#39 Mistral-7B-Instruct-v0.32.2M
#40 Qwen2.5-32B-Instruct1.9M
#41 gemma-4-12B-it1.9M
#42 Qwen3.5-27B1.9M
#43 SmolLM2-135M-Instruct1.8M
#44 SmolLM2-135M1.8M
#45 GLM-4.7-Flash1.8M
#46 Mistral-7B-Instruct-v0.21.7M
#47 Qwen2.5-Coder-14B-Instruct1.7M
#48 Qwen2.5-14B-Instruct1.6M
#49 Qwen3.5-35B-A3B1.6M
#50 Qwen3-30B-A3B1.5M
#ModelDownloads / 30dFormatNeedsSmallest device that holds itDecode
1Qwen3-0.6B
Qwen · 752M params
29.7M
1.7K likes
Q4_K_M2.22 GBNVIDIA GeForce GTX 1650
4 GB
~65 tok/s
2Qwen3-VL-8B-Instruct
Qwen · 8.8B params
14.6M
1.2K likes
Q4_K_M7.04 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~70 tok/s
3gemma-4-26B-A4B-it
Google · 26B params
13M
1.6K likes
Q4_K_M16.8 GBAMD Radeon RX 7900 XT
20 GB
~234 tok/s
4Qwen3-8B
Qwen · 8.2B params
10.7M
2.1K likes
Q4_K_M7.04 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~70 tok/s
5gemma-4-31B-it
Google · 31B params
9.9M
4K likes
Q4_K_M22.2 GBNVIDIA GeForce RTX 4090
24 GB
~33 tok/s
6Qwen3.5-9B
Qwen · 9.7B params
9M
2.1K likes
Q4_K_M7.05 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~85 tok/s
7Qwen2.5-0.5B-Instruct
Qwen · 494M params
8.7M
653 likes
Q4_K_M1.31 GBNVIDIA GeForce GTX 1650
4 GB
~203 tok/s
8Qwen2.5-7B-Instruct
Qwen · 7.6B params
8.4M
2.4K likes
Q4_K_M5.95 GBNVIDIA GeForce RTX 2060
6 GB
~47 tok/s
9Qwen3.5-4B
Qwen · 4.7B params
7.8M
1K likes
Q4_K_M3.98 GBNVIDIA GeForce GTX 1650
4 GB
~32 tok/s
10Qwen3-4B
Qwen · 4.0B params
7.8M
720 likes
Q4_K_M4.51 GBNVIDIA GeForce RTX 2060
6 GB
~62 tok/s
11Qwen2.5-1.5B-Instruct
Qwen · 1.5B params
7.4M
863 likes
Q4_K_M2.15 GBNVIDIA GeForce GTX 1650
4 GB
~72 tok/s
12Qwen3.8-27B
Qwen · 28B params
6.9M
16.9K likes
Q4_K_M18.6 GBAMD Radeon RX 7900 XT
20 GB
~42 tok/s
13gpt-oss-20b
OpenAI · 21B params
6.6M
5.1K likes
Q4_K_M13.8 GBAMD Radeon™ RX 9070 XT
16 GB
~199 tok/s
14Qwen2.5-VL-7B-Instruct
Qwen · 8.3B params
5.8M
1.7K likes
Q4_K_M6.36 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~85 tok/s
15GLM-5.3-Flash
zai-org · 321B params
5.4M
2.7K likes
Q4_K_M198.3 GBAMD Instinct MI325X Accelerator
256 GB
~495 tok/s
16Qwen3.5-2B
Qwen · 2.3B params
4.9M
423 likes
Q4_K_M2.32 GBNVIDIA GeForce GTX 1650
4 GB
~71 tok/s
17dolphin-2.9.1-yi-1.5-34b
dphn · 34B params
4.7M
69 likes
Q4_K_M23.5 GBNVIDIA GeForce RTX 4090
24 GB
~31 tok/s
18DeepSeek-V4-Flash-0731
DeepSeek · 304B params
4.5M
4K likes
Q4_K_M187.5 GBAMD Instinct MI300X Accelerator
192 GB
~367 tok/s
19gpt-oss-120b
OpenAI · 117B params
4.5M
5.3K likes
Q4_K_M71.9 GBNVIDIA RTX PRO 5000 Blackwell 72GB
72 GB
~257 tok/s
20gemma-4-E4B-it
Google · 8.0B params
4.4M
1.6K likes
Q4_K_M5.86 GBNVIDIA GeForce RTX 2060
6 GB
~76 tok/s
21Qwen2.5-3B-Instruct
Qwen · 3.1B params
4.2M
587 likes
Q4_K_M3.21 GBNVIDIA GeForce GTX 1650
4 GB
~39 tok/s
22Qwen3-32B
Qwen · 33B params
4M
754 likes
Q4_K_M22.7 GBNVIDIA GeForce RTX 4090
24 GB
~32 tok/s
23Qwen3-4B-Instruct-2507
Qwen · 4.0B params
3.7M
989 likes
Q4_K_M4.72 GBNVIDIA GeForce RTX 2060
6 GB
~62 tok/s
24DeepSeek-V3.2
DeepSeek · 685B params
3.6M
1.5K likes
Q4_K_M422.2 GBAMD Instinct MI455X
432 GB
~842 tok/s
25Qwen3-VL-4B-Instruct
Qwen · 4.4B params
3.5M
485 likes
Q4_K_M4.72 GBNVIDIA GeForce RTX 2060
6 GB
~62 tok/s
26pythia-160m
EleutherAI · 213M params
3.5M
46 likes
Q4_K_M1.01 GBNVIDIA GeForce GTX 1650
4 GB
~496 tok/s
27NVIDIA-Nemotron-3-Nano-4B-BF16
NVIDIA · 4.0B params
3.4M
126 likes
Q4_K_M3.46 GBNVIDIA GeForce GTX 1650
4 GB
~33 tok/s
28Qwen3.6-35B-A3B
Qwen · 36B params
3.3M
2.9K likes
Q4_K_M23.1 GBNVIDIA GeForce RTX 4090
24 GB
~370 tok/s
29Qwen3-1.7B
Qwen · 2.0B params
3.1M
562 likes
Q4_K_M3.02 GBNVIDIA GeForce GTX 1650
4 GB
~43 tok/s
30gemma-4-E2B-it
Google · 5.1B params
3M
999 likes
Q4_K_M4.02 GBNVIDIA GeForce RTX 2060
6 GB
~200 tok/s
31OTel-2.0-LLM-31B-IT
farbodtavakkoli · 31B params
3M
22 likes
Q4_K_M22.2 GBNVIDIA GeForce RTX 4090
24 GB
~33 tok/s
32Qwen3-VL-2B-Instruct
Qwen · 2.1B params
2.8M
478 likes
Q4_K_M3.05 GBNVIDIA GeForce GTX 1650
4 GB
~43 tok/s
33Qwen3-14B
Qwen · 15B params
2.7M
487 likes
Q4_K_M11.1 GBNVIDIA GeForce RTX 5070
12 GB
~46 tok/s
34Qwen3.5-0.8B
Qwen · 873M params
2.6M
742 likes
Q4_K_M1.46 GBNVIDIA GeForce GTX 1650
4 GB
~155 tok/s
35JiRackUltra_1b
CMSManhattan · 1.8B params
2.5M
0 likes
Q4_K_M2.15 GBNVIDIA GeForce GTX 1650
4 GB
~72 tok/s
36Qwen3.6-27B
Qwen · 28B params
2.5M
2.3K likes
Q4_K_M18.6 GBAMD Radeon RX 7900 XT
20 GB
~42 tok/s
37Qwen2.5-VL-3B-Instruct
Qwen · 3.8B params
2.4M
709 likes
Q4_K_M3.41 GBNVIDIA GeForce GTX 1650
4 GB
~39 tok/s
38Qwen2.5-Coder-7B-Instruct
Qwen · 7.6B params
2.3M
813 likes
Q4_K_M5.95 GBNVIDIA GeForce RTX 2060
6 GB
~47 tok/s
39Mistral-7B-Instruct-v0.3
Mistral AI · 7.2B params
2.2M
3.7K likes
Q4_K_M6.25 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~77 tok/s
40Qwen2.5-32B-Instruct
Qwen · 33B params
1.9M
362 likes
Q4_K_M22.8 GBNVIDIA GeForce RTX 4090
24 GB
~32 tok/s
41gemma-4-12B-it
Google · 12B params
1.9M
1.6K likes
Q4_K_M9.54 GBNVIDIA GeForce RTX 3080
10 GB
~63 tok/s
42Qwen3.5-27B
Qwen · 28B params
1.9M
1.1K likes
Q4_K_M18.6 GBAMD Radeon RX 7900 XT
20 GB
~42 tok/s
43SmolLM2-135M-Instruct
HuggingFaceTB · 135M params
1.8M
430 likes
Q4_K_M1.09 GBNVIDIA GeForce GTX 1650
4 GB
~316 tok/s
44SmolLM2-135M
HuggingFaceTB · 135M params
1.8M
242 likes
Q4_K_M1.09 GBNVIDIA GeForce GTX 1650
4 GB
~316 tok/s
45GLM-4.7-Flash
zai-org · 31B params
1.8M
1.9K likes
Q4_K_M20.4 GBNVIDIA GeForce RTX 4090
24 GB
~35 tok/s
46Mistral-7B-Instruct-v0.2
Mistral AI · 7.2B params
1.7M
3.2K likes
Q4_K_M6.24 GBNVIDIA GeForce RTX 3070 Ti
8 GB
~77 tok/s
47Qwen2.5-Coder-14B-Instruct
Qwen · 15B params
1.7M
192 likes
Q4_K_M11.4 GBNVIDIA GeForce RTX 5070
12 GB
~45 tok/s
48Qwen2.5-14B-Instruct
Qwen · 15B params
1.6M
372 likes
Q4_K_M11.4 GBNVIDIA GeForce RTX 5070
12 GB
~45 tok/s
49Qwen3.5-35B-A3B
Qwen · 36B params
1.6M
1.5K likes
Q4_K_M23.1 GBNVIDIA GeForce RTX 4090
24 GB
~370 tok/s
50Qwen3-30B-A3B
Qwen · 31B params
1.5M
953 likes
Q4_K_M20.2 GBNVIDIA GeForce RTX 4090
24 GB
~251 tok/s
How the requirement column is computed

Each row is a real run of the same engine the calculator uses: weights at the quantization shown, the KV cache at 8,192 tokens in F16, and the runtime reserve — with offload switched off, so “holds it” means fully resident rather than technically loadable.

Devices are ordered by published capacity, not by price. This site has no price data, and ranking by an invented one would be exactly the kind of unsourced figure it refuses everywhere else.

Looking at it from the other side — what does my card run? — every hardware page answers that for one device.

llmbottleneck
catalogue 2026-10-03models 327devices 135