LLMBOTTLENECK.COM
Method / source hierarchy / formulas

Show your work

The engine separates facts from constructions and constructions from guesses. Every result carries the weakest evidence level of the inputs that actually determine it, and names which input that was.

In short
  • Memory = the weight file + the KV cache (the conversation) + a runtime reserve. The answer is whether all three fit in the device’s own memory.
  • Speed is mostly decided by memory bandwidth: every token reads the active weights again. The rate is fitted to published benchmarks per backend and format.
  • Every figure is labelled published data, estimate or unvalidated — and a result carries the weakest label of the inputs that decide it.
  • Nothing is filled in: a layer type or a figure with no published source is refused, not guessed.

Three evidence levels

published data

Checked against an official immutable source: a pinned config.json, a manufacturer specification, or a published GGUF file.

estimate

A transparent derivation from verified inputs, such as KV geometry or a fitted decode estimate whose error is published.

unvalidated

A useful but unmeasured assumption. It stays visible and can never quietly become a published fact.

What the labels mean
published data
Read from a published source — a file's byte count, a model's configuration or a manufacturer's specification — or exact arithmetic on such values. Not a measurement on a machine.
estimate
Calculated from sourced inputs by a stated method; an estimate, not a measurement.
unvalidated
Rests on at least one assumption that has not been measured, such as the runtime reserve.
reconstructed size
Weight size reconstructed from the pinned architecture, because no published file exists.
size range
Only a lower and an upper bound are claimed for this weight size.

Weight files: three tiers

The competitor problem this fixes is real: a calculator that only answers for models with a published GGUF answers for a sixth of a catalogue, and one that multiplies parameters by a rounded bits-per-weight is wrong by tens of percent on exactly the small models people run first. This engine does neither.

  1. Verified. A published GGUF exists for the pair. Its byte count is used as reported, with a link to the pinned revision.
  2. Derived. No file exists, so one is reconstructed. Every tensor's shape comes from the pinned configuration, every tensor's encoded size comes from the exact ggml block layout, and the type each tensor gets comes from llama.cpp's mixture rule. A reconciliation gap larger than 0.1% is carried into the published range.
  3. Bounded. The architecture cannot be reconstructed to within 2% of the parameter count Hugging Face publishes for the same revision. Nothing is claimed except the range the block layouts allow, and the interface shows a range.
weight file size
tensor bytes = ceil(elements / block elements) × block bytes, padded to 32
file bytes   = Σ tensor bytes + tokenizer header allowance

A model only reaches the derived tier after the reconstruction reconciles with the parameter count published for the same pinned revision. That gate is the whole safeguard: an architecture we cannot reproduce is an architecture whose tensor mix we have no right to price.

Block layouts, from the pinned source

TypeElements per blockBytes per blockBits per weight
F321432.0000
F161216.0000
Q4_032184.5000
Q5_032225.5000
Q8_032348.5000
Q2_K256842.6250
Q3_K2561103.4375
Q4_K2561444.5000
Q5_K2561765.5000
Q6_K2562106.5625

All read from ggml at commit 62acc89c26c6. Bits per weight is derived from the block geometry, never assumed.

The mixture rule

llama.cpp does not quantize every tensor the same way. Normalisation tensors and the expert router stay in 32-bit floats; the output projection is raised; and on the _M mixtures the value projection and the down projection are raised on a fixed subset of layers:

llama.cpp mixture rule
raised(layer) = layer < L/8  or  layer ≥ 7L/8  or  (layer − L/8) mod 3 == 2

That predicate is transcribed from the pinned quantization routine and then confirmed independently against the tensor types recorded in real GGUF headers in the catalogue. For Q4_K_M, Q5_K_M, Q5_0, Q6_K, Q8_0 and FP16 the whole mapping is header-confirmed. For Q4_0, Q3_K_M and Q2_K it is transcribed only, which is why those formats carry a visibly larger published error.

Memory

device requirement
total = weight file + KV cache + runtime reserve
KV    = 2 × layers × KV heads × head dimension × context × batch × encoded bytes

Grouped-query, multi-query, latent attention and per-layer sliding windows all use their official configuration geometry rather than a correction factor. Recurrent and linear-attention layers hold a constant state, sized from its own formula instead of the KV-head one, and sparse-attention (DSA) layers add the cache their indexer keeps. A layer type with no formula of its own is refused rather than sized as if it were conventional attention. The runtime reserve is still an uncalibrated prior, and every result that includes it says so.

Decode

decode time per token
seconds/token = c₀ + (active weight bytes + KV bytes) / (MBU × bandwidth)
tokens/second = 1 / seconds/token

Active weight bytes, not resident bytes. A routed model reads its shared tensors plus the experts a token actually visits, so a 30B file with 8 of 128 experts per token decodes at a fraction of what its size suggests. At batch sizes above one, more experts get touched, and the expected number follows directly from each token routing independently:

experts touched per token
experts touched = E × (1 − (1 − k/E)^batch)

MBU and c₀ are fitted per backend and format. Where a format has no fit of its own, the nearest fitted format on the same backend is used and the result says so. That substitution is small for a measured reason: on Metal, the only backend with three independent fits, utilisation moves by 1.2% across the whole range from 4.5 to 16 bits per weight.

Prefill and time to first token

first-token floors
compute floor = prompt tokens × 2 × active parameters / peak FLOPS
memory floor  = active weight bytes / bandwidth
reported      = [max(floors), max(compute floor / MFU, memory floor / MBU)]

Utilisation is solved from one published benchmark rather than copied from a competitor: 9.6% of the published FP32 peak on amd-radeon-ai-pro-r9700. One anchor cannot support a point estimate, so a range is reported instead. The reference peak is an FP32 vector figure; hardware with dedicated matrix engines beats it, which is why a compute-bound result here is the slow end of what to expect. Where a vendor publishes no compute peak, only the memory roof is priced and the answer is labelled a lower bound.

Offload

offloaded token time
token time = c₀ + GPU bytes/(MBU × GPU BW) + host bytes/effective RAM BW + KV time

llama.cpp offloads whole layers, so the split is quantised to layers rather than to a byte fraction: a half-offloaded layer does not exist. Offloaded layers execute on the host; they are not modelled as a fictional per-token copy across PCIe. Effective system bandwidth is a declared input, not a measurement of your machine, and results say so.

Diagnosis policy

The order is deterministic: no fit, offload limited, capacity limited, bandwidth limited, balanced. Balanced is only reachable after real alternatives have been evaluated, and an upgrade is only recommended when the smallest step that clears a 20% gain exists. A recommendation names one resource, quantifies the change, and says what not to buy.

A bandwidth verdict does not inherit the runtime reserve's evidence level. The reserve is identical on both sides of a device comparison, so it cancels; marking that verdict a hypothesis because of a term that does not affect it would be its own kind of dishonesty.

Citing this

LLM Bottleneck, “Methodology”. https://llmbottleneck.com/methodology. Model architecture from official Hugging Face repositories at the pinned revisions listed on each model page; ggml block layouts from ggml-org/llama.cpp at commit 62acc89c26c6.

Read the measured error of all of this →

llmbottleneck
catalogue 2026-10-03models 327devices 135