LLMBOTTLENECK.COM
Public scorecard / no cherry-picking

Accuracy, with the scope attached

A small honest score is more useful than a universal-looking percentage. Each card says exactly what was compared, and the format this engine handles worst is on the same table as the ones it handles best.

In short
  • Published file sizes are used exactly as published: 0% median error over 318 files.
  • Rebuilt file sizes, for formats nobody published, land within a median 0.017% of the real file over 224 cases — worst case 26.3%, explained below.
  • Speed estimates are off by a median 2.6%, with 100% of 39 held-out benchmarks within 20%. Measured on 39 held-out benchmarks of one model, Llama 2 7B, 87% of them on Apple hardware. It is not evidence about a mixture-of-experts model, a long context, a batch of concurrent requests, or a device generation the corpus does not contain.
  • Not measured: the memory a real run peaks at. Every answer on the site says so.
Two different questions

This page measures how close the engine’s calculations land. Whether the inputs are right is a separate question: every figure on a model page names the publisher’s own file it was read from, at a pinned revision, so it can be checked without taking this site’s word for it. A perfect formula fed a wrong layer count is still a wrong answer, so both have to be published.

Four different claims, and only two are measured

“How much memory does this need?” is four questions stacked on each other, and a score for one of them is not a score for the next. This page measures the bottom two. The top two are calculated under stated assumptions and are labelled that way everywhere they appear.

What this scorecard covers and what it does not
The claimEvidence behind itScored here
Bytes of the published fileThe publisher’s own file listing, at a pinned revisionYes — 318 files
Bytes when no file exists, rebuilt from architectureReplayed against every published file and compared byte for byteYes — 224 cases
Resident memory: weights, KV cache and runtime reserveCalculated from the pinned geometry. The reserve is an uncalibrated prior, and the result says so on its faceNo corpus
Peak memory a run actually touchesDepends on the runtime’s allocator, its buffers, fragmentation and the workload — none of which a file size impliesNot claimed at all

The distance between the second row and the fourth is the honest gap in this product. Closing it needs runs on physical hardware, not a better formula, and until those exist no number here is presented as one.

Reconstructed weight files vs published files

estimate

320 of the 327 catalogued models can be sized, including the ones with no published GGUF; the other 7 publish neither a file nor a usable parameter count, and are refused rather than guessed. To find out how much that costs in accuracy, the reconstruction is replayed against every published file we do have and compared byte for byte.

Cases224published files
Median error0.017%of file size
Within 1%70%of cases
Inside the band96%range we publish
FormatCasesMedianWorstInside band
FP16220.005%2.47%100%
Q5_0180.007%9.48%100%
Q6_K280.010%13.76%93%
Q5_K_M290.011%11.07%93%
Q8_0420.012%25.91%98%
Q4_0140.014%1.66%93%
Q4_K_M400.031%26.30%93%
Q3_K_M151.712%7.64%100%
Q2_K161.758%3.21%100%

The worst column is where the honesty lives. Q2_K and Q3_K_M are worse than the rest: their mixture was wrong until published headers corrected it during the last data pass, and llama.cpp still varies those two formats with model properties this rule does not carry. Every large individual outlier is a small model, where the embedding matrices are a big share of the file and one unknowable layout choice — whether the published file ties its output tensor — moves the size by tens of percent. That case is not hidden behind a point estimate: the range widens and the interface shows the range.

Excluding those tie-ambiguous models, the remaining 110 cases have a median error of 0.008% and a worst case of 26.30%.

  • The corpus is the catalogued models whose own publisher released a GGUF. It spans eleven publishers, which is what makes it evidence that the reconstruction generalises rather than fitting one publisher's conventions, but it is still not a random sample of the catalogue.
  • The GGUF metadata header is modelled from the tokenizer vocabulary at forty bytes per entry, rounded outward from the 39.3 bytes per entry measured on the one artifact family whose headers are parsed.
  • Q2_K and Q3_K_M carry a visibly larger error. Their mixture was corrected from published headers during this pass, but llama.cpp also varies those two formats with model properties this rule does not carry, notably on routed models.
  • A model that does not reconcile with the published parameter count is never sized this way; it falls back to the block-geometry bounds.
Published-file retrieval
published data
Cases318
Median error0%

This checks that a published file's byte count is used exactly as Hugging Face reports it. A 0% error here is retrieval correctness, not a prediction victory, and it says nothing about total runtime memory.

Decode calibration
estimate
LOOCV cases39
Median error2.64%
Worst case11.83%
Within 20%100%

Measured on 39 held-out benchmarks of one model, Llama 2 7B, 87% of them on Apple hardware. It is not evidence about a mixture-of-experts model, a long context, a batch of concurrent requests, or a device generation the corpus does not contain.

Decode-speed calibration coverage and error, by hardware vendor and by backend
SegmentCasesShare of corpusMedian errorWorst case
Apple vendor3487%2.89%10.59%
AMD vendor38%1.83%2.19%
NVIDIA vendor25%11.20%11.83%
Metal backend3487%2.89%10.59%
Vulkan backend38%1.83%2.19%
CUDA backend25%11.20%11.83%
  • Leave-one-out validation covers one Llama 2 7B corpus only; it does not establish universal accuracy.
  • The corpus contains 34 Apple, 2 NVIDIA, and 3 AMD rows across Metal, CUDA, and Vulkan.
  • The active weight bytes are the explicit GGUF artifact bytes; kvBytes is held at zero for this decode-only corpus.
  • No MoE active-weight inference, offload correction, batch correction, or tensor-parallel correction is evaluated.
Decode speed, next to the other published figure

Memory is arithmetic, and every serious tool lands close. Decode speed is where they separate — and it is the number the nearest comparable calculator publishes as its own weakest. Both figures below are each site’s own self-assessment on its own corpus, read from its own accuracy page on 1 September 2026.

Published decode-speed error, this site and apxml.com
Published figureThis siteapxml.com
Median error, decode speed2.6%15.1%
Geometric mean factor error1.04×1.22×
Worst case11.8%0.30× (a 70% miss)
Cases scored3950
ValidationLeave-one-out, so no case scores its own calibrationNot stated

What this site’s column was measured on: Measured on 39 held-out benchmarks of one model, Llama 2 7B, 87% of them on Apple hardware. It is not evidence about a mixture-of-experts model, a long context, a batch of concurrent requests, or a device generation the corpus does not contain.

Read this narrowly. Different corpora, different case counts, and each number chosen by the site reporting it — so this is not a controlled comparison and nothing here says their engine is worse. What it does show is that the figure they publish as their weakest is the figure this site holds itself to hardest, and that the cases here are scored leave-one-out: every case is predicted by a calibration fitted without it, which is the only way a self-reported error rate means anything at all.

Time to first token is a range, not a number

estimate

Prefill utilisation is solved from a single published benchmark: 9.6% of the 191 TFLOPS dense matrix peak its vendor publishes, from 1024 prompt tokens at 3009.02 tok/s on one device.

That is one measurement, on one backend, so the site does not present a point estimate. It shows the physical roofline floor and the utilisation-adjusted figure as a range.

Read the range, not its upper end. The benchmark behind the utilisation is a routed mixture-of-experts model on a Vulkan build, and routing prefill through eight experts per token issues many small matrix multiplications where a dense model issues few large ones. A dense model on a CUDA or ROCm build keeps the matrix units far busier than 9.6% and should be expected to land nearer the floor than the calibrated figure. Both ends are stated for exactly that reason; a single number here would be the one that is wrong.

Prefill is compute bound and a runtime dispatches its matrix multiplications to the tensor or matrix units, so the roof has to be priced from those units and not from the headline vector figure — on an RTX 3090 the two differ by a factor of two, on an RTX 6000 Ada by a factor of four. The catalogue now carries a dense matrix peak, quoted from the vendor's own specification table, for 21 of 135 devices, and every time to first token on the site is priced against it.

What remains withheld. The other 114 devices get no time to first token, and the calculator says why on each one. Two reasons account for all of them. Some vendors publish matrix throughput only with structured sparsity — NVIDIA quotes H200 at 1,979 FP16 teraFLOPS footnoted “With sparsity”, which is exactly twice the dense rate and describes a workload with half its weights pruned; halving it here would be this site's arithmetic rather than the vendor's publication, so the estimate is withheld. Others, Apple silicon among them, publish no matrix throughput for the GPU at all. The memory roof alone is a few milliseconds; publishing it as a “lower bound” would be dressing a vacuous claim as an answer.

Open the benchmark this rests on ↗

MoE: published observation, not a model

estimate

Qwen3-30B-A3B Q4_K_M on Radeon AI PRO R9700, llama.cpp Vulkan build 8233: 183.47 tok/s at 128 and 171.30 tok/s at 4,096.

The benchmark did not publish a GGUF hash or repository link. Filename and exact byte size match the pinned official artifact; identity is therefore founded, not verified. This single device/runtime observation is not extrapolated.

Open the primary benchmark and raw commands ↗

What is not claimed

  • No measured accuracy for total runtime VRAM. The runtime reserve is an uncalibrated prior and every result says so on its face.
  • No measured accuracy for offloaded decode, multi-user serving, or routed decode speed. Those are models resting on published physics, not fitted to a corpus of their own.
  • No time to first token on the 114 devices whose vendor publishes no dense matrix throughput — either because the only figure offered is sparsity-doubled, or because none is offered at all. No dense figure is ever inferred from a sparse one.
  • No accuracy claim for the 116 of 327 catalogued models whose architecture the engine cannot reconstruct. Those are sized with block-geometry bounds and shown as a range on purpose.

Every table on this page is generated by a script that fails the build when a threshold is missed, so a regression cannot quietly widen these numbers. The method behind them is on the methodology page.

llmbottleneck
catalogue 2026-10-03models 327devices 135