Two quantizations were being added together
Discovery resolved a file's quantization by searching for a known name followed by a separator, and what follows the separator is usually the rest of the same name: Q6_K_L contains Q6_K followed by an underscore. A repository publishing both had the two summed into one artifact, recording Trinity-Large-Thinking Q6_K as 687 GB when it weighs 343, which put it 53% outside its own published band. The same defect applied to Q4_0_4_8 against Q4_0, despite the comment claiming to guard against exactly that.
The token is now read whole and then recognised, never searched for by prefix. Worst published-size error fell from 52.7% to 26.3%.
Nine files sit outside the derived band, and are listed
Three repositories ship mixtures that are not the standard one for the name on the file: a “Q4_K_M” of 6.34 effective bits per weight is not a stock Q4_K_M. In all three the parameter count reconciles and the Q8_0 file lands exactly, so the geometry is right and the mixture is not. Those nine are named in the test suite rather than absorbed by widening every band on the site, and they do not reach a reader: when a published file exists the engine serves its actual byte count. Band coverage of the derived tier is 95.7%.
Time to first token now covers 16 devices instead of 3
Prefill runs its matrix multiplications on the tensor or matrix units, so pricing its compute roof from the headline vector figure is wrong by the ratio between the two: a factor of two on an RTX 3090, four on an RTX 6000 Ada. Until now the estimate was published only where a vendor's own label said the figure was a vector one, which meant three AMD cards and nothing else. The catalogue now carries a dense matrix peak quoted from the vendor's specification table for sixteen devices, and the calibration is re-solved against the benchmarked card's own matrix peak of 191 TFLOPS, giving 9.58% utilisation where the vector solve gave 38.30%. Every published time to first token moved with it.
A sparsity figure is never halved into a dense one
NVIDIA quotes H200 at 1,979 FP16 teraFLOPS footnoted “With sparsity”, and the RTX PRO 6000 Blackwell at 4,000 AI TOPS footnoted “Theoretical FP4 TOPS using sparsity”. Those numbers are exactly twice the dense rate and describe a workload with half its weights pruned, which a dense transformer is not. Rather than divide by two and present the result as the vendor's, those devices carry no matrix peak and no time to first token, and each says so in its own caveat. The same applies to Apple silicon, which publishes no matrix throughput for the GPU at all.
Published headers corrected the low-bit mixtures
Q2_K was over-estimating every routed model by a 7.99% median with a systematic +7.28% bias. The transcribed rule raised the value and down projections on every layer; the published headers say Q2_K raises neither and holds both uniformly at Q3_K, while Q3_K_M raises them from Q4_K to Q5_K on the usual subset. Reading the real headers instead of inferring brought Q2_K to a 1.44% median with the bias gone.
A reconstruction now carries its own reconciliation gap
A model reproducing its published parameter count to 1.6% was being refused a reconstruction entirely and priced with block bounds instead, which made the most downloaded model on Hugging Face show a range twice as wide as it needed. Such a model is now reconstructed, and its own gap is added to the published range rather than ignored, so the range says how well that particular architecture is understood. Gemma 4 12B moved from a 5.1 to 12 GB range to 7.0 to 8.1 GB.
Two data bugs found by widening the corpus
A multimodal projector file was being read as a model's weights, recording a 397B model as weighing 0.9 GB. And one publisher's safetensors index reports a parameter count its own published bytes contradict by five orders of magnitude. Both are now refused: a projector is not weights, and a parameter count implying more than nine bytes per parameter is not usable.
The lower bound admits what a conversion drops
A published parameter count describes the whole checkpoint, but a GGUF holds only what llama.cpp runs: a multimodal vision tower and a multi-token-prediction head go into a separate file or nowhere. Ten published files sat below a bound that assumed otherwise. The bound now allows for it, measured at the largest observed shortfall.
Catalogue rebuilt: 160 models, 24 publishers
Discovery was ranking each publisher by downloads, which is dominated by older repositories and hid every recent release. It now pins the current flagships explicitly and covers publishers it had never queried. Gemma 4, Mistral Small 3.2, Magistral, Devstral, Xiaomi MiMo, Tencent Hy4, MiniMax M2.7, GLM-5.3 Flash, Nemotron 3.5, granite 4.2, Ling 3.0, Ornith 1.5 and ERNIE 4.5 are all in the snapshot for the first time.
Why Llama is not here
Every meta-llama and older google repository on Hugging Face sits behind a manual licence gate, and a gated repository does not serve its config.json: the request returns 401. Architecture is only accepted from the model's own publisher, so the Llama family cannot be catalogued without a licence acceptance this project does not hold. Gemma 4 is a different case and is included, because Google published it ungated.
Official GGUF coverage more than tripled
A discovery pass now looks for a GGUF repository owned by the same publisher as each catalogued model, so the exactly-known weight tier grew from 108 artifacts across 16 models to 335 across 54, spanning eleven publishers instead of one. Community requantizations are still never admitted: a file is only used when its own publisher released it.
The size model held up outside its training family
Until now the reconstruction had only ever been scored against Qwen files. Replayed against the enlarged corpus it lands at a 0.012% median across eleven publishers, which is the first evidence that it generalises rather than fitting one publisher's conventions.
Two config shapes that were being rejected
A published max_window_layers of zero means every layer is windowed; the validator was treating it as invalid and dropping the model. Hunyuan publishes moe_intermediate_size as one entry per layer, and MiMo publishes moe_layer_freq as a per-layer mask. Uniform per-layer arrays are now collapsed to the value they all carry, non-uniform ones are refused rather than averaged, and the per-layer mask is kept in its own field instead of being discarded.
Hardware: RTX 50 series, DGX Spark, M4 and M5 Pro
Added the RTX 5070 Ti, RTX 5070 and RTX 5060 Ti 16GB from NVIDIA's comparison specifications, DGX Spark from its product page, and the M4, M4 Pro and M5 Pro from Apple's own announcements, each field carrying the sentence that states it.
Two devices deliberately left out
AMD's product pages refused every automated request during this pass, so the RX 9070 non-XT, RX 9060 XT and RX 7800 XT wait for a specification we can quote from AMD directly. The Xiaomi AI Cube, announced on 24 August, publishes 80GB of unified memory and 1.22TB/s of near-memory computing bandwidth: that figure describes compute adjacent to memory on its accelerator, not the streaming bandwidth this engine's decode model consumes, and Xiaomi has published no specification page for it. Entering it would produce confident tokens-per-second figures with nothing behind them.
Every model now answers, in three tiers
The calculator previously offered only the sixteen models with a published GGUF. Weight files are now sized on a tier ladder: a published artifact where one exists, a reconstruction from the pinned architecture and exact ggml block layouts where none does, and block-geometry bounds where the architecture cannot be reconstructed. The tier travels with every number, and the reconstruction's measured error is published per format on the accuracy page.
Routed models are priced on the bytes a token reads
MoE decode was withheld entirely. It is now computed from the shared tensors plus the experts a token actually visits, with the expected number of distinct experts derived from independent routing at batch sizes above one. Capacity still requires the whole file resident, which is why a routed model can be fast and still not fit.
Time to first token is published as a range
Prefill utilisation is solved from one published benchmark rather than inherited from a competitor, so a point estimate would overstate what is known. The physical roofline floor and the utilisation-adjusted figure are shown together, and a device whose vendor publishes no compute peak gets a memory-roof lower bound instead of silence.
Offload and concurrency reach the interface
Layer offload, batch, concurrent users, runtime and declared system bandwidth are now inputs rather than a note saying they are unavailable. Offload is quantised to whole layers because that is what llama.cpp does. Free-plan ceilings clamp the result and name the capability that was clamped.
A speed verdict no longer inherits the runtime reserve
The uncalibrated runtime reserve is identical on both sides of a device comparison, so it cancels. Marking every bandwidth verdict a hypothesis because of a term that does not affect it was its own kind of dishonesty. The reserve still governs the evidence level of capacity verdicts, where it genuinely binds, and every result now names its weakest input.
Latest-model snapshot refreshed
Refreshed four model entries with pinned official configs: DeepSeek-V4-Pro-0813, GLM-5.3, LFM2-24B-A2B and LFM2-2.6B-Longevity. The catalog remains at 100 models; recency did not override source or coverage invariants.
DSA false precision blocked
DeepSeek V4 and GLM-MoE-DSA return unsupported for KV memory. Their official implementations maintain compressed, expanded or indexer state that the conventional KV-head formula does not represent.
Radeon AI PRO R9700 source added
Added AMD's official 32 GB workstation GPU specification: 640 GB/s bandwidth, 47.8 FP32 TFLOPS and 300 W. Its published llama.cpp results are used as a founded observation and, since this change, as the single anchor behind the prefill calibration.
RTX PRO 6000 Blackwell source added
Added the Workstation Edition from NVIDIA's official specification: 96 GB GDDR7 ECC, 1,792 GB/s, 125 single-precision TFLOPS and 600 W. Bus width remains blank because NVIDIA does not publish it on that source.
GGUF representation invariant
Corrected 20 artifacts where a monolithic GGUF and its equivalent multipart publication had been summed together. Added ingestion and runtime invariants that reject mixed representations.
Qwen3.8 MTP boundary
Recorded that the official config declares one MTP layer while the pinned standard Transformers model ignores MTP keys. That mismatch is why the model does not reconcile and is sized with bounds rather than a reconstruction.