LLMBOTTLENECK.COM
Every gap on the record

What this catalogue does not hold

A gap is allowed to exist. A gap nobody wrote down is not. Every release that shows real demand is either in the catalogue or listed below with a reason read from the publisher's own configuration file — and the check that enforces that runs on every build and once a day.

Held224/395

Repositories that show demand and are pinned to a revision in the catalogue. Demand means a commercial market sells tokens from it, or it ranks in Hugging Face’s own trending, download or like lists.

This year114/159

Releases from 2026. The rest are all accounted for below; an unaccounted one fails the build rather than sitting unnoticed.

Queued0

Models the engine can size that nobody has ingested yet. Each carries the date it was detected and fails the build after 14 days. Without an expiry a list like this is where an awkward model goes to be forgotten.

Four states, not two

“Supported” hides more than it says. A model can be detected and not sized, sized and not benchmarked. Each row below says which.

What each coverage state promises
StateWhat it promises
HeldArchitecture read from the publisher’s own config.json at a pinned revision, and the engine can produce a memory figure from it.
QueuedThe geometry is a shape this engine already models. Nothing is claimed about it until it is ingested.
UnsupportedNaming it would be easy and answering about it would be wrong. The reason is architectural and is stated per model.
Answered by its baseA requantization, fine-tune, merge or adapter whose geometry is the base model’s. Counting it separately would inflate the catalogue without answering a new question.

Cannot be held, and why

published data

Every reason below describes something about the repository, not about this project’s convenience. “Not got round to it” is not a reason; that is what the queue above is for. Each is re-checked when the catalogue is refreshed, because a reason that has stopped being true is a gap again.

nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-FP8

A re-export of a checkpoint already held

An FP8 re-export of a Nemotron 3 checkpoint, carrying the architecture of the BF16 repository it derives from.

nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

Architecture this engine does not model

Corrected 2026-09-13 (audit M03): the earlier reason said the config publishes no layer count, which was wrong — llm_config.num_hidden_layers is 52, with a hybrid_override_pattern. What is missing is this engine's support: the geometry is nested under llm_config of a multimodal (audio/vision) wrapper that the ingest does not read, and the vision and audio encoders are not modelled. It stays excluded as a limitation of this engine, not of the publisher's data.

zai-org/GLM-5.3-BF16

A re-export of a checkpoint already held

A dtype re-export of zai-org/GLM-5.3, which the catalogue holds. The publisher tags no base_model, so the derivative rule cannot see the relationship, but the geometry is the same file's.

inclusionAI/Ling-3.0-flash-VL

A modality this engine does not size

A vision-language model whose config nests its text geometry under text_config and requires the publisher's own modelling code (auto_map to modeling_bailing_moe_v3_vl). This engine sizes text decoding only; it models neither the vision tower, the projector nor visual tokens, so a figure from it would answer a narrower question than the model poses.

https://huggingface.co/inclusionAI/Ling-3.0-flash-VL/resolve/main/config.json read 2026-09-10; model_type bailing_moe_v3_vl

IFM/K2-Horizon-7B-Uno

The publisher states no architecture

The repository serves no config.json at all: the hub answers 'Entry not found' for it on the default revision. There is no published geometry to pin, and inventing one is the failure this catalogue exists to avoid.

https://huggingface.co/IFM/K2-Horizon-7B-Uno/resolve/main/config.json returned 'Entry not found' on 2026-09-10

chenyumo/moziAI-35B-A3B-MOE-MTP

A fine-tune of a model already held, with the same geometry

An uncensored fine-tune of ornith-ai/Ornith-1.5-35B-A3B, which the catalogue holds: the model card names that base, and the repository publishes only a Q4_K_M GGUF of the result. The publisher tags no base_model, so the derivative rule cannot see the relationship, but the text geometry in its config.json is Ornith-1.5-35B-A3B's field for field — 40 layers, hidden 2048, 256 experts with 8 per token, and the same linear/full-attention layer_types — so every memory and speed answer is the base model's.

https://huggingface.co/chenyumo/moziAI-35B-A3B-MOE-MTP/resolve/f9d39d75a378598ae603b28928249c314c067ffb/config.json and README.md (Base Model: Ornith-1.5-35B-A3B) read 2026-09-11; model_type qwen3_5_moe

nex-agi/Nex-N2.5-Max

Architecture this engine does not model

A DeepSeek-V4 derivative (DeepseekV4ForCausalLM, 1.6T parameters) whose layer_types name heavily_compressed_attention alongside compressed_sparse_attention. Rechecked 2026-09-20: the compression rates themselves are published (compress_rates gives 128 for the heavily compressed layers and 4 for the sparse ones, and compress_ratios repeats them per layer), so the missing piece is not the rate but the topology. The config names no kv_source_layer_ids, so which layer writes the 128-to-1 cache and which layers read it is unstated, and the cross-layer formula this engine uses for DeepSeek-V4.1 cannot be applied without inventing that mapping. It is recorded as a limitation of this engine's reading, not of the publisher's data.

https://huggingface.co/nex-agi/Nex-N2.5-Max/resolve/e45b5be454592dcd1ffa82640904747a6b11b7fa/config.json read 2026-09-11; model_type deepseek_v4

Agnuxo/CAJAL-4B

The publisher states no architecture

The repository publishes only GGUF files (f16, q4_k_m, q8_0) and no config.json, so there is no published geometry to pin; its model card names no base model either. A file size alone says nothing about the KV cache.

https://huggingface.co/Agnuxo/CAJAL-4B/resolve/2ab404efe5aa75c4fc7e0d15de771923faf3e653/config.json returned HTTP 404 on 2026-09-11

meta-llama/Llama-4-Maverick-17B-128E-Instruct

restricted_official

Meta's repository is licence-gated: the hub answers 401 for its config.json, so no geometry can be pinned from it. The catalogue holds unsloth/Llama-4-Maverick-17B-128E-Instruct instead, an ungated copy of the same release, recorded as a mirror in data/model-identity.json and not byte-compared with Meta's file. The gap is the licence gate, not the architecture.

CohereLabs/North-Small-Translate-1.0

restricted_official

Access-gated (gated: auto) and licensed cc-by-nc-4.0: the hub answers 401 for its config.json, so there is no published geometry to pin, and the licence would not permit the commercial use this tool's readers are sizing hardware for. Its pipeline tag is translation rather than text generation. Recorded so that the absence is a decision rather than an oversight; no dimensions are guessed for it.

CMSManhattan/JiRackUltra_14b

derivative_of_held_model

A redistribution of DeepSeek-R1-Distill-Qwen-14B, which the catalogue holds. The publisher tags no base_model, so the derivative rule cannot see the relationship, but its config.json is that model's geometry field for field — Qwen2ForCausalLM, hidden 5120, 48 layers, 40 attention and 8 key-value heads, intermediate 13824, vocabulary 152064 — and the model card names a DeepSeek R1 14B base. The repository ships GGUF quantizations and a single model.safetensors with no weight index, so the downloads that surface it are file pulls of those quantizations; at 2 likes against a million downloads the count measures traffic rather than adoption. Every memory and speed answer is the base model's.

CMSManhattan/JiRackDeltaNet_27b

derivative_of_held_model

A redistribution of Qwen/Qwen3.8-27B, which the catalogue holds. The publisher tags no base_model, so the derivative rule cannot see the relationship, but its config.json (revision fc9d90d5531e4a70c2a699435fa0bc90065f6d96) is the held model's text_config field for field — hidden 5120, 64 layers, 24 attention and 4 key-value heads, head_dim 256, intermediate 17408, vocabulary 248320, full attention every fourth layer, 262,144-token maximum, and the same linear-attention geometry — differing only in declaring the text-only Qwen3_5ForCausalLM head. The model card calls it ternary, but nothing in the config says so and its published Q4_K_M file is 16.8 GB, the size of an ordinary 27B quantization; the card's own benchmark table is Qwen3.8-27B's, credited to that repository. As with the same publisher's JiRackUltra_14b, the repository ships GGUF quantizations and one unindexed model.safetensors, so the download count measures file pulls. Every memory and speed answer is Qwen3.8-27B's.

ReliquaryForge/qwen3-4b-base-dapo-v4

derivative_of_held_model

A DAPO fine-tune of Qwen/Qwen3-4B-Base, which the catalogue now holds (added 2026-09-20). The publisher tags no base_model, so the derivative rule cannot see the relationship, but the geometry is the base's field for field — Qwen3ForCausalLM, hidden 2560, 36 layers, 32 attention and 8 key-value heads, head_dim 128, intermediate 9728, vocabulary 151936, 32,768-token maximum and full attention throughout — so every memory and speed answer is Qwen3-4B-Base's.

Qwen/Qwen-72B

Architecture this engine does not model

The first-generation Qwen architecture (QWenLMHeadModel, model_type qwen, auto_map to configuration_qwen.py). Its config publishes intermediate_size 49152, which is not the feed-forward width: the publisher's modelling code halves it, so the SwiGLU branches are 24576 wide. It also publishes no num_key_value_heads and states its head width as kv_channels rather than head_dim. Normalizing it with the shared reader would state a feed-forward width twice the published weights', which is the kind of confident wrong answer this catalogue exists to avoid. A 2023 release, held back by the reader rather than by the repository.

apple/OpenELM-1_1B

Architecture this engine does not model

OpenELM scales its layers individually: ffn_multipliers gives a different feed-forward width for each of the 28 layers and num_query_heads / num_kv_heads give a different head count for each. The catalogue represents one width and one head count per model, so there is no honest single value to record, and collapsing the arrays would misstate every layer but one. A limitation of this catalogue's shape, not of Apple's data, and the licence (apple-amlr) is its own restriction.

Cactus-Compute/needle3

Architecture this engine does not model

Reviewed again on 2026-10-04, when the conventional cache learned separate key and value row widths (for MiMo-V2). That was the limitation recorded here, and it no longer is one. Three things in the config at revision c7c415a3d1b3 still have no reading in this engine. The key width is published as qk_head_dim 48 with no head_dim, and the reader takes the key width from head_dim. Four of the twenty layers are global, but they are listed as extras.global_attention_layers [4, 9, 14, 19], nested where the reader does not look, so the stack would be priced as if every layer were windowed. And the config states two windows, sliding_window 1024 and kv_window 256, without saying which bounds the cache. Cactus-Compute/needle2, whose config states a single head_dim, one window and no global layers, is held. A limitation of this reader, not of the publisher's data.

IFM/K2-Horizon-MoVA-36B-A4B

Architecture this engine does not model

Every sparse layer routes its value projection through 64 additional experts (mova_num_experts), which the publisher's K2HorizonMoVAAttention builds as 64 Linear(hidden_size -> num_key_value_heads x head_dim) modules per layer. That is 64 x 2560 x 1024 weights on each of the 45 sparse layers, about 7.5B parameters, none of which this engine's reconstruction accounts for: it would understate a 36B checkpoint by roughly a fifth while reporting the feed-forward mixture correctly. The cache is unaffected — the mixed value keeps the published key-value head shape — so the gap is in weights, not in context. The dense K2-Horizon-0.9B and 7B checkpoints are held.

paradigma-inc/limite-1b-violetto

Architecture this engine does not model

Read at revision b1f3d572ccac. The config publishes a hybrid stack by listing which layers are global: global_layers [3, 7, 11, …, 47], twelve of the 48, with global_window -1 against a sliding_window of 1024 for the rest. This reader has no field for that list, so the engine would price all 48 layers as windowed and leave out the twelve that keep the whole context, the part of the cache that grows. The weights do not reconcile either: reconstructing the published geometry gives 1,190,585,600 parameters against the 1,035,253,888 in the repository's safetensors, 15% more, because the model's own components (value embeddings on sixteen layers, MUDD dense connections on two, per-head attention gates) are built by the publisher's modeling_limite.py in ways the reconstruction does not follow. An answer from this engine would be wrong on both terms.

TaichuAI/ZDTaichu5.0-9B

Architecture this engine does not model

Read at revision bc125a9819a4. A vision-language model whose text geometry is nested under llm_config (a Qwen3_5ForCausalLM of 32 layers and hidden size 4096) inside a wrapper class loaded from the publisher's own modeling.py, beside a C-RADIO vision tower. The ingest reads geometry from the top level or from text_config, not from llm_config, and this engine models neither the vision tower, the projector nor visual tokens. Its 9,794,197,512 published parameters include that tower. The same limitation as the Nemotron Omni wrapper recorded above: of this engine, not of the publisher's data.

Hardware: what is catalogued, and what is not

A device is catalogued only when its manufacturer publishes the memory and the memory bandwidth the engine needs. Where that figure is missing the card is listed here, not approximated from a neighbour with a similar name. Reviewed 2026-09-23.

Not catalogued: a required figure is not published · 3

  • M1 (base) · Apple, accessible

    Apple has never published a memory bandwidth for M1; later pages give it only as a ratio ("50 percent more than M1", "nearly 6x … M1"), which is not a published figure. Checked on apple.com and support.apple.com specification pages.

  • Ryzen AI Max PRO 490 and 485 · PCs and mini-PCs

    AMD's series table gives these two parts the same 192 GB unified memory as the Max+ PRO 495 and publishes no separate graphics-memory ceiling, memory speed or bandwidth for either. They differ from the 495 in cores, clocks and graphics compute units, none of which this engine prices. Catalogued separately they would be three names for one set of memory figures, which would overstate how much of the market this catalogue has actually read.

  • Radeon RX 5700 XT · Radeon, installed base

    AMD's product page for the RX 5700 XT returns 404 (checked 2026-09-23 at /en/products/graphics/desktops/radeon/5000-series/amd-radeon-rx-5700-xt.html). A third-party figure would be acceptable under the 2026-09-23 rule; the entry waits until one is read and quoted.

Not yet read · 1

  • GTX 1050 Ti 4 GB · GeForce, installed base

    4 GB holds only models under about 3B parameters. Its figures are available from the same third-party source as the other GTX 10 cards; it is next in the queue ordered by Valve's survey share, not excluded.

Not modelled by this engine · 2

  • GB200 / GB300 platforms · Data centre, NVIDIA

    Rack-scale platforms with CPU and GPU memory tiers; the engine sizes one device type per split and does not model a Grace-Blackwell rack. The B200 and B300 GPUs are catalogued on their own.

  • CPU-only inference with system RAM · CPU and APU

    The engine sizes accelerators and unified-memory machines; a CPU mode with channels and DIMM bandwidth is not implemented.

Catalogued · 42

What this page does not claim

  • That the catalogue holds every model ever published. It deliberately does not: it holds the ones people are running now, and says how that is decided.
  • That a release is detected within any guaranteed interval. The check runs daily and records when each gap was first seen, so the real interval can be measured rather than asserted.
  • That a held model has been benchmarked. Holding it means its geometry is pinned and sizeable; measured speed is a separate and much smaller corpus, published on the accuracy page.

Generated from the same audit the build runs, on 2026-10-03.

llmbottleneck
catalogue 2026-10-03models 327devices 135