LLMBOTTLENECK.COM
Corrections are product features

Changelog

Every change to a source, an assumption or the engine, newest first — corrections included.

2026-10-04

1 entry

Any Hugging Face repository, thirteen new models, and a correction to MiMo's cache

Paste a Hugging Face link on the new /hf page and the site answers for repositories it does not catalogue. Most downloads are fine-tunes and GGUF files of a model rather than the model itself, so the page follows the base_model tag the repository's own publisher wrote, through a GGUF repository to the fine-tune it quantizes if need be, and compares its config.json with the catalogued base's on every field that decides memory. When they match, the base model's answers are the repository's, and the files it publishes are listed with their own sizes and the memory each needs.

When nothing was published to compare, as with a GGUF-only repository, the page says the answer rests on the publisher's tag alone; when the geometry differs or no base is declared, it refuses and says why. A base is never read out of a repository's name. The same answer is available as GET /v1/resolve, and publishers can put a memory badge on their model card.

The derivatives the daily coverage audit has already compared have their own indexed pages and are linked from their base model. A correction: Xiaomi's MiMo-V2 models were sized as if every layer attended over a 128-token window. Their configuration marks nine of the 48 layers of the Flash models, and ten of the 70 of the Pro models, as full attention, and in the Flash models gives the windowed layers eight key-value heads rather than four; both are in the publisher's modelling code at every pinned revision, and the layer split is in its model card.

Their cache was therefore understated, by little at short context and by a great deal at long context: MiMo-V2.6-Flash at 262,144 tokens needs 6.07 GB of cache, not 0.02 GB. The engine now reads the per-layer pattern and sizes key and value rows separately. Weight sizes do not change.

Thirteen models join the catalogue, taking it to 327: Aleph Alpha's Kolibri-1, released on 2 October and reconstructed to within 288,000 of its 78.1 billion published parameters; Yandex's AliceAI-Foundation-80B-A3B-Base; the MOPD upgrades of MiMo-V2.6 Pro and Flash; LLM-jp 4.1 at 8B, 33B and 32B-A3B; and six fine-tunes with demand of their own, among them MiMo-V2.6-Distill-Qwen-9B, Apple's LensVLM-9B and H Company's Holo4. Two releases that cannot be held are recorded with the reason, and a third's reason is brought up to date. The downloadable dataset is regenerated.

2026-10-01

1 entry

One default context, one snapshot id, and a command that works

The calculator and the widget now open at 8,192 tokens of context, the figure every model, device and pair page is sized at; they opened at 4,096, so the same pair read 6.43 GB on the home page and 7.04 GB on its own page. The snapshot id is now derived from a canonical text of each data file: the site's bundler and Node read a few long fractions one unit apart in the last place and so produced different ids from identical files, which meant the check that the dataset matched the site could not fail. No figure changes; the id does.

The dataset manifest records the snapshot its rows were last verified against, and a rebuild that changes no row no longer publishes a second copy. The command-line tool is served by the site as one file to run with Node, because the npx command the integrations page printed named a package that was not on the registry. The buying guide names a card for a model only when at least 5% of the card stays free: it suggested a 72 GB card for a model needing 71.9 GB.

The index of memory sizes names the largest popular model that fits rather than the absolute largest. Decode rates read the same everywhere, as a rounded estimate with a tilde. The privacy page no longer says both that no newsletter exists and how to subscribe to it, and it now describes benchmark submissions, whose submitter digest is keyed with a deployment secret instead of a constant in the source.

2026-09-29

1 entry

Which GPU to buy for local LLMs

A buying guide built from the same engine as every other page. By memory size — 8, 12, 16, 24, 32 GB and more — the fastest card and the most common one in Steam's hardware survey, how many models each holds entirely and how fast it runs the same 8B model; and the other way round, the smallest card that holds each of the most downloaded models. Within one memory size every card runs the same models, so "fastest" is simply the most memory bandwidth. It prints no prices, because none can be sourced; the store links open a search.

2026-09-27

1 entry

What runs on your amount of VRAM

A new page for each amount of memory a catalogued device has — 8, 12, 16, 24 GB and more — lists every model that fits entirely at Q4_K_M with 8,192 tokens of context, the largest one, and how fast each runs. Fit depends only on memory, so the list holds for every card of that size; speed depends on bandwidth, so it is quoted for one named card, the most common of that size in Valve's hardware survey, with the others listed beside it. There is no new arithmetic: the rows are the same engine call as each device's own page.

2026-09-23

1 entry

45 more devices, including laptops, and a labelled tier for third-party figures

The device catalogue grows from 90 to 135, chosen by Valve's published hardware survey rather than by hand: the RTX 20, GTX 16 and GTX 10 desktop cards, the RTX 3050 8 GB, the Radeon RX 6600 to 6950 XT, RX 7600 and 7700 XT, and 22 RTX 30, 40 and 50 laptop GPUs. Together they take the share of surveyed graphics processors the site can answer for from about 40% to 73%. Every Radeon figure and the RTX 50 laptop bandwidths are read from AMD's and NVIDIA's own pages.

NVIDIA no longer publishes the memory bandwidth of its older and laptop parts, and has no page left for the GTX 10 series, so those figures come from the TechPowerUp GPU Database: they are labelled estimates, quote the page they came from, and each device says so in its caveats. The catalogue's validator accepts a third-party figure only on those terms. Four 2026 releases with demand join the model catalogue: IFM's K2-Horizon-3.7B and K2-Horizon-375B-A23B, Altworld's Hemmingway-1 and CMSManhattan's JiRackUltra_1b.

2026-09-22

2 entries

Rent a GPU that runs it: live cloud prices, per model and per card

When your card cannot hold a model, or only holds it at a degraded format, its page now shows the rented machines that do, with this site's speed estimate and cost per million tokens for each, and the price Vast.ai and RunPod quote right now, read from their own public APIs. It opens with what your own card can still do — the offload speed, free — so renting is weighed against the free option rather than presented as the only one. One machine is marked recommended, by a rule printed next to it and used everywhere on the site: the lowest cost per million tokens among machines that answer at 20 tokens a second or more.

Beside it sit the cheapest, and the cheapest machine at least 1.6× faster. Every model page ranks every rentable configuration, one card to eight, at Q4_K_M, Q8_0 or FP16, and puts the per-token price from OpenRouter and Novita AI beside it, with the ratio that decides between them for one person. Where the catalogue holds the exact GGUF file, the page gives the container image and start command that load it, and a four-step guide says how renting works for anyone who has not done it — including the part that costs people money, which is forgetting to stop the machine.

Every GPU a provider rents out has a panel with its price by provider and count, and a calculator that answers whether buying the card or renting it is cheaper at your hours and electricity price. The new Cloud GPUs page compares every rentable card side by side with what it runs. Some of the provider buttons are referral links: Vast.ai, RunPod and Novita pay this site a share of what a new customer spends.

The order is always the price and never the commission, OpenRouter — which pays nothing — is listed first whenever it is cheaper, and the terms of every programme, and why others were left out, are published on the Cloud GPUs page.

Xiaomi MiMo-V2.6 Pro and Flash, and the packed checkpoint the engine had to read

Xiaomi's MiMo-V2.6 open weights join the catalogue, taking it to 310: MiMo-V2.6-Pro-RL (1.02T total and 42B active parameters) and MiMo-V2.6-Flash-RL (309B total and 15B active), both sparse MoE with a 1M-token context, read from the publisher's own config.json at the release revision and credited to Xiaomi. The checkpoints store their expert tensors as MXFP4 U8 blocks — two four-bit values per element — so Hugging Face's element total (524.1B and 159.4B) counts storage words rather than logical parameters. The engine now reads the declared store_dtype, withholds that word count and sizes the pinned architecture against the card's own logical total, carrying the reconciliation gap into the range instead of hiding it. Both card figures are recorded in data/model-parameter-facts.json with their source hashed.

2026-09-21

1 entry

New API plans: Pro and Business, yearly billing, and a one-step checkout

The API now has two paid plans. Pro ($19 a month or $190 a year) lets you show results without the LLM Bottleneck credit, runs the widget unbranded on your domain, and allows 10,000 calls a day. Business ($79 a month or $790 a year) runs the widget unbranded on up to ten domains with your own button in it — a store search or an affiliate link, filled in with the model and graphics card the reader chose — at 100,000 calls a day, with priority support.

Enterprise is available by contract. The free plan keeps every endpoint at full fidelity, 200 calls a day, with a visible credit wherever results are shown publicly. Subscribing takes one button: Stripe's checkout collects the payment (card, Apple Pay, Google Pay or Link) and, for businesses, a VAT number, and your API key is waiting on the next page with the plan already active.

Existing Pro+ subscribers keep their price. Also new: GET /v1/changes lists the models and devices added or updated since a date; widget domains are managed from Usage & keys; TypeScript and Python clients can be downloaded from /v1/clients; and the model and pairing pages show the exact API request and widget tag for the configuration you are looking at.

2026-09-20

1 entry

Seventeen new models, two new devices, and more accurate page dates

Seventeen models join the catalogue, taking it to 308, among them NVIDIA Nemotron 3 Ultra 550B, Mistral Large 3 675B, Llama 4 Maverick, Step-3.5-Flash, sarvam-105b, GigaChat 3.5 and EXAONE 4.0. Most were absent because the reader could not read what their publishers had published: Nemotron 3 Ultra writes its layer count as one entry per layer, Mistral Large 3 ships its own params.json instead of a Transformers config, and StepFun's layer list continues through three multi-token-prediction heads. Each is now read, and each reading is checked against the publisher's own weight index.

Every model that stays out is recorded with a reason that can be checked. NVIDIA Rubin and the AMD Ryzen AI Max+ PRO 495 join the hardware catalogue from their makers' own figures, Rubin marked as announced. A sitemap lastmod now records when a page's own content last changed, rather than stamping every page with the date a catalogue was rebuilt.

The key-issuance limit of forty a day is actually counted. The content security policy covers scripts and every other fetch. The decode-speed accuracy figure now says, wherever it appears, that it was measured on one model and mostly on Apple hardware.

The downloadable dataset is regenerated from the current catalogue.

2026-09-14

1 entry

Billing, multi-GPU sizing and new data-centre cards

Plan changes now happen on your existing subscription: upgrades charge the prorated difference straight away, downgrades credit the unused time on the next invoice, and a checkout can never turn into a second purchase. Your daily allowance belongs to the account, so rotating or adding keys no longer changes it. Tensor-parallel sizing now places key-value heads the way runtimes do, which makes very large models across many GPUs more accurate on the cluster, target and recommendation tools.

A link to a specific Mac configuration now opens exactly that machine. The widget offers the same configurations as the calculator. The A6000, L40S, B200 and B300 join the catalogue, along with the RTX 4060, the 40 SUPER series and the 30-series cards that were missing.

Search ignores punctuation, so "Llama 3.3" finds the model, and the privacy page lists every service that processes data, Stripe included.

2026-09-13

2 entries

The Macs and Arc cards people already own, and every source checked again

Fifteen devices join the catalogue, taking it to 76. Seven are earlier Apple silicon chips the catalogue skipped — M1 Pro and M1 Max, M2, M2 Pro and M2 Max, M3 and M3 Pro — each read from Apple's own technical specifications page, which prints the memory bandwidth and the memory each machine was sold with. The base M1 stays out: Apple never published its bandwidth, only ratios to it on later pages, and a figure divided out of a marketing ratio is not a published one.

Six are Intel cards — Arc Pro B70 and B65 (both 32 GB at 608 GB/s), Arc B570, and the Arc A770 16 GB, A770 8 GB and A750 — read from Intel's specification pages; the two A770s are separate devices because Intel publishes 560 GB/s for the 16 GB card and 512 GB/s for the 8 GB one. Two are AMD Instinct modules, the MI350X and MI455X, recorded per module rather than as the racks they are sold in. Complete machines are now listed on the page of the chip they are built on, with what the maker says limits the GPU: Framework Desktop and HP Z2 Mini G1a (up to 96 GB assignable), Dell Pro Max with GB10 and ASUS Ascent GX10.

The Arc Pro B50's Intel page had moved; its figures were re-read at the new address and are unchanged. Every numeric specification with a source was then checked against that source: 257 of 260 values were found on the page they cite, one more by reading a table whose extracted layout is ambiguous, and the remaining two sit behind a document viewer and a cookie-gated table and are listed as not checked. That check now runs weekly.

The coverage page lists, for each hardware priority, whether it is catalogued or why not. The rules for reusing figures, the dataset and third-party material are now stated once and repeated identically on the terms, the data page and the citation page. The API returns a request id with every response, logs latency and outcome per route, and has a status endpoint; TypeScript and Python clients are generated from the OpenAPI document and tested against the real handlers.

The derivative check no longer trusts a base-model tag: a fine-tune is answered by its base only when its configuration has the same geometry, which caught pruned mixtures of experts and untied embeddings.

Your exact machine, unified memory done right, and a steadier API

This release makes answers match the machine you actually have. A Mac was sized at the largest configuration Apple sells: a 36 GB M4 Max was answered as a 128 GB one at 546 GB/s instead of its own 410 GB/s. The calculator now asks for your configuration, the API accepts memory_gb and hardware_variant, and an answer sized without them says so.

Unified memory was treated as two pools, so a 72B model on a 32 GB Mac was offered an offload into 32 GB of system RAM the machine does not have; offload now applies only where a second pool exists, and a configuration that does not run gets no speed figure. /v1/cluster counted a card's memory in GiB while everything else counted GB, which turned one no-fit into a fit; every surface now uses the same budget, and the difference between installed and counted memory is reported instead of assumed. One checkpoint's parameter count had counted packed 4-bit storage words, sizing a 27B model as a 12 GB FP16 file; re-reading its metadata at the same pinned revision corrected it, and the engine now refuses that fingerprint. LoRA estimates omitted the adapter weights and applied a Llama layout to every architecture; both are fixed and fine-tuning is labelled beta.

The API now validates one scenario schema on every route, refuses unknown fields, echoes what it applied, reserves a whole batch atomically and never charges for invalid input. Forty-three models were added — including Llama 3.2 and 3.3, Gemma 2 and 3, Qwen2.5 small sizes, Qwen2.5-VL and Coder, Qwen3.5 0.8B and 2B — and eleven devices, including the RTX 5060, 5050, RX 7900 XT, 7800 XT and 7600 XT. The dataset now has a row with a stated status for every model and device pair, and its files are named by content hash so a cited version never changes.

2026-09-12

2 entries

The whole table, as a file anybody can check

Everything this engine knows about what runs where is now one download: 11,000 rows pairing 241 models with 50 devices at 8,192 tokens, each carrying the best format that fits, the memory it needs, and — the part that makes it worth publishing rather than recomputing — the evidence level behind both the weight size and the device bandwidth. CSV and JSON, versioned, with a SHA-256 for each file so a copy that has been edited on its way to you does not match. The licence draws the line this site has always drawn in words: the compilation and the computed columns are CC BY 4.0, and the facts underneath — a manufacturer's published bandwidth, a publisher's configuration — are not ours to license and are not claimed.

One condition is stated rather than assumed: the evidence columns are data, not decoration, and republishing a reconstructed figure as a measured one misrepresents the work rather than merely reusing it. In the same pass, sarvam-30b joined the catalogue, taking it to 248. It had been queued eight hours earlier by the gap detector and the queue slot said the engine could hold it; reading its configuration confirmed it — nineteen layers, sixty-four query heads over four key/value heads, a 128-expert mixture routing six plus one shared, and no sliding window — so it was ingested rather than left to expire.

Bing gets two things it can use without an account: a verification token read from the environment, so the tag appears the moment somebody signs in and pastes it rather than sitting empty and invalid, and IndexNow, whose key file proves control of the host and lets a catalogue refresh announce itself instead of waiting to be noticed. Finally, the rule that decides which pairing pages ask to be indexed can now read Search Console. It only ever promotes: a page with impressions is indexable whatever the rule guessed, and a page with none is never demoted for it, because a page marked noindex is never shown and so can never earn the impressions that would rescue it.

That trap is the reason the rule stayed categorical until there was data to read.

The cards that were missing because nobody published one number

The RTX 3060 12GB and the RTX 4060 Ti 16GB are most of the cheap 12 and 16 GB of video memory in people's machines, and this catalogue refused both of them for ten days. The reason was real: NVIDIA publishes their memory type and interface width and stopped publishing their bandwidth, and taking the figure from a third-party database would have put an unsourced number under thousands of generated pages. The refusal was right about the database and wrong about the arithmetic.

Bandwidth is the interface width times the data rate over eight; NVIDIA publishes the width, and the company that builds the board publishes the rate for the exact card in the box. Both pages are now recorded, the formula is stored with them, and the result carries a new evidence level, derived — 192-bit at 15 Gbps is 360 GB/s for the 3060, 128-bit at 18 Gbps is 288 GB/s for the 4060 Ti. The validator will not accept a derived figure unless its evidence states arithmetic that reproduces the stored number, and a test recomputes every one of them.

Intel enters the catalogue at the same time, which required a fourth vendor in a schema that had three: the Arc B580 at 456 GB/s, and the Arc Pro B60 24GB and B50 16GB, which are currently the cheapest way to put that much accelerator memory in a workstation. Three AMD parts join them — the RX 9060 XT 16GB, the RX 9070 and the Instinct MI355X — and they nearly did not, because amd.com times out on a plain fetch and renders its specification tables into collapsed accordions that a text scrape reads as empty. They were read from the vendor's own rendered table in a browser instead, which is the same standard as every other row here.

The catalogue goes from 42 devices to 50. One correction fell out of the same pass: Apple sells the M6 at two memory bandwidths under one name, 153 GB/s on the two lower Mac mini models and 170 GB/s on the highest, and this catalogue held only the maximum — so an owner of the cheaper machine was reading a decode estimate from 11% more bandwidth than their computer has. Both are now recorded, and the pages say which is which.

2026-09-11

2 entries

DeepSeek-V4.1-Flash, once its cache was read the way it is built

Refused a day ago because the per-layer formula would have overstated its cache roughly tenfold, DeepSeek-V4.1-Flash is now in the catalogue. The engine reads two fields it used to ignore — kv_source_layer_ids and compress_ratios — and prices only the four layers that write the long-range cache, each entry a 512-wide FP4 latent with E4M3 scales per 16 channels plus a 128-wide FP4 index key, with every layer’s own 128-token FP8 window on top. That arithmetic gives 890 bytes per token of global cache, the figure DeepSeek’s technical report publishes, and a test now holds it there.

The formats are the publisher’s, from the report and the reference inference code rather than config.json, so the KV format chosen in the calculator does not change them and the answer says so. Two things are counted conservatively and stated: the 196B Engram parameters are included in the weights although DeepSeek’s own serving stack prefetches them from host memory, and weight sizes are bounded rather than derived, because no GGUF exists and the architecture does not reconcile to a published count. Six other August and September releases join it — MiniCPM5-2B, Nex-N2.5-Pro, Nex-N2.5-mini, Spark-X2.5-4B, NeoHorse-1-4B and LFM2.5-VL-3B — each pinned to the revision read on 11 September, taking the catalogue to 247.

Motif-3 and K2-Horizon-MoVA stay out: the first interleaves windowed layers on a pattern the engine would misread, and the second routes its values through experts the engine does not model.

Two hundred and forty models in one alphabetical run

The model picker listed the catalogue by repository name, so Qwen3.8 sat forty rows from Qwen3.6, Llama appeared under a mirror’s namespace and nothing said which release was a week old. It now groups by lab and by family, with both headings pinned while their rows scroll; the labs readers come looking for are listed first and community fine-tunes last, and within a lab the newest line leads. A side rail filters to one lab, to the most downloaded, or to the last six weeks’ releases with their dates.

Search ignores punctuation, so “qwen 3.6” and “v4.1 flash” find what they mean, best match first. Sizes carry a decimal below 100B, because a row reading 15B under a model named Qwen3-14B looked like an error even though it is the files’ own count. The device picker is grouped by vendor and product line the same way.

2026-09-10

9 entries

An answer for eight, read as the answer for thirty-two

A request for a batch of 32 across 32 concurrent users returned HTTP 200 carrying batch 8 and concurrency 8. The figures were correct for eight, and a caller who did not inspect a nested gates object had no reason to doubt they were the answer to what they asked. The cause was a commercial ceiling being enforced inside the physical engine, and the same confusion ran through the pricing page, whose comparison table promised batch above eight free while the engine quietly lowered it.

Batch and concurrency are not features: they enter the calculation as one multiplier on the KV cache and cost this server nothing, so they have left the plan table entirely. What remains are engineering bounds — 1,024 each, 8,192 sequence slots — that refuse by name and return a 422. Nothing below them is substituted, ever.

The rule now has no third branch: the scenario is calculated as asked, or it is refused, and there is no path that answers a different question quietly.

A context ceiling nobody had measured, attributed to the publisher

The page for Qwen3-8B on an RTX 4090 said 32,768 tokens was the publisher’s own published maximum. Qwen’s pinned config.json says 40,960. The page had walked the context up in powers of two and concluded, from the next doubling exceeding the model’s limit, that the last one it tried was that limit — a search that never evaluates 40,960 cannot conclude anything about 40,960, and could only ever report a power of two whatever the truth was.

Bisection over the admissible range finds the real wall: 40,960 on a 4090, which genuinely fits at 23.2 GB of 24, and 9,147 on a 3070, where the limit is memory rather than the model and no doubling could have expressed it. Four things that were one number are now four fields: the measured ceiling, the publisher’s own maximum, the nearest round setting for someone planning in recognisable figures, and what the limit actually is. The publisher’s maximum is asserted only when that exact value was evaluated and fit; where a model declares none, the page says the figure is where our search stopped, which is a statement about us rather than about the model.

The calculator had the same defect from the other side — its control only accepted powers of two, so 40,960 was not expressible on the day the pair page was corrected to report it.

Sharing a configuration restored a different one

The permalink carried the model, format, device, context, KV and some cluster fields. The analysis request carried all of those plus runtime, offload, host RAM and custom device memory. They were built independently, so a shared link opened a different question from the one that had been asked and presented the answer to it with no indication anything had changed.

Both are now projections of one versioned scenario, so they cannot drift apart again, and a round trip through a link is asserted under test. Fields equal to a default are written out rather than omitted, because a default that later changes would otherwise silently rewrite every link ever shared. A parameter that cannot be read is named back to the caller instead of being replaced by a default — an unknown interconnect used to be dropped and answered around.

The economic assumptions travel in the link too: a cost per million tokens is a claim, and a link that dropped it restored a page making different ones.

Published file sizes are not measurements, and the site said they were

A heading reading MEASURED, NOT ESTIMATED sat above a paragraph about published file sizes, and the verdict carried a badge saying “88% of it measured”. That figure is a share of the memory total whose bytes come from published sources, not a probability that the answer is right, and it invites exactly the second reading. Nothing on this site is measured on a running machine unless it comes from the benchmark corpus, which is labelled separately.

The badge now says what it counts. The pair tables called their rows “published formats” while including sizes reconstructed from architecture, and now say how many of each: five of the nine formats for Qwen3-8B are sized from a published file and the rest are rebuilt. Decode and first-token figures say they are calibrated estimates rather than runs on that card, and the homepage no longer promises how fast a model will actually go.

The gap-detector that was never run, and a model it was right to refuse

The audit that finds catalogue gaps existed and ran only when somebody chose to run it: the check script did not include it and nothing scheduled it, so “we track new releases” described a script rather than anything that happened. It is now part of the check and runs daily. Its first honest run named nine 2026 releases unaccounted for, all now on the record with reasons read from their own configs.

One of them is DeepSeek-V4.1-Flash, published the same day, and it is recorded as unsupported rather than added: its config shares a single KV cache across layers through kv_source_layer_ids — four producing layers for forty decoder layers — with per-layer compression and a 128-token window over the compressed cache, so the per-layer formula would overstate the cache by roughly an order of magnitude. It also declares engram layers of 384 million embeddings each, which no transformer formula contains. Supporting it on the generic path would have produced a confident and badly wrong number.

Three models the engine genuinely can hold are queued instead of excluded, with a detection date and a fourteen-day target that fails the build — an exclusions file with no expiry is where an awkward model goes to be forgotten.

A verdict before the question, and a highlight nobody could see

On a phone the results panel came first and the form second, so the page opened on “IT FITS” about a configuration the visitor had not chosen, with the model picker roughly 2,900px down. It is now 575px, before any verdict, with the refinements below the answer; the desktop two-column layout is unchanged. The model picker had a related defect of its own: the arrow keys moved a highlight with no aria-activedescendant behind it, so a screen reader announced the first option and then nothing, and the active option was never scrolled into view, so past the eighth model the same was true for everyone.

The control knew which option was current and told nobody. Both are fixed, along with Home and End, a spoken match count, 44px rows, and a mark that distinguishes the option in effect from the one being arrowed over.

Which catalogue answered, and a door that was standing open

A computed answer was cached for a day in shared caches, justified by it being a pure function of files that ship with the build. That reasoning holds right up to the moment the build changes: an entry that is still fresh never revalidates, and an ETag is only consulted after expiry, so a catalogue refresh could go unnoticed by caches for a day with nothing in the response saying which catalogue had produced the figures. Every answer now names its snapshot — engine version, both catalogue versions, their sizes — and the calculator sends that id, so a refresh changes the cache key outright rather than hoping a cache asks.

Asking for a snapshot this build cannot produce returns the current figures and a warning saying so. Separately, the unkeyed calculator endpoint had no rate limit at all, which made it the cheapest way to use this engine in volume; it now has one, described precisely as what it is rather than as a defence it is not.

Blocked and de-indexed at the same time, which achieves neither

robots.txt disallowed /embed while /embed also served a noindex. Those are opposite instruments: a crawler refused the page never reads the directive on it, so the block prevented the snippet and not the listing, and a link from any customer’s site could still put the URL in an index. Crawling, indexing and the sitemap are now one declared decision per route with the reason attached, and a test asserts nothing is ever in both states. /usage leaves the sitemap because it shows nothing without an API key. priority and changeFrequency are gone, Google having said it does not use them, and lastmod no longer claims the privacy policy was revised because a GPU was added. In the same pass the pair pages stopped selecting the thirty smallest devices in the catalogue under a comment claiming they spanned its range — which meant the pages answering “which card does run this” were never generated for the large ones.

The site was not on the internet

The domain had been registered, its nameservers were the platform’s and DNS resolved to the platform’s anycast, but it was attached to no project — so the edge held a domain with no deployment to serve from it, and every visitor got a platform 404. Every canonical URL on the site pointed at it, and the widget snippet handed to shops loaded its script from it. It is deployed and verified from outside: home, robots, sitemap, a pair page, the JSON endpoint, the embed frame, widget.js and the OpenAPI document all answer, www redirects to the apex, and the widget renders inside a page served from a different origin.

Deploying immediately found something local testing could not: self-service API keys are stored in a JSON file, and a serverless filesystem is read-only, so issuance answers with a refusal. The endpoint was already failing honestly rather than handing out a key that would never authenticate; what was wrong is that the pricing page went on listing it as built. It now says it is closed, and why, before anyone fills in the form.

2026-09-05

4 entries

A wrong layer count, priced in gigabytes

The cross-check could say that another calculator's table disagrees with a model publisher's own file, which is true and, on its own, boring: nobody chooses hardware from a head count. What people act on is the KV cache, and those are exactly the fields that multiply it — so each disagreement is now run through this site's own attention engine twice, over the publisher's geometry and over the one the other page publishes, and reported as the difference in gigabytes at a length someone would really use. Only the disputed fields are substituted; layer types, sliding windows and latent attention come from the publisher's file for both runs, which is the reading most favourable to the other site.

The worst case is a model whose cache lands 2.13× and 7.11 GB too large at 32,768 tokens — the difference between one accelerator and two. This also replaced a piece of arithmetic that was quietly wrong here: the old ratio multiplied layers by KV heads by head dimension, which is right for a plain grouped-query model and wrong for everything the catalogue has been filling up with. On a latent-attention model it reported a 0.01× error where the real cost is nothing at all, and the page now says so in those words.

The analysis endpoint answers on GET, so the answer can be cached

An analysis here is a pure function of the pinned catalogue: the same configuration returns identical bytes until the catalogue is refreshed and redeployed. It was only reachable by POST, which no cache may reuse, so every embedded widget impression on somebody else's product page paid a server round trip for an answer already computed thousands of times. It now also answers on GET, using the permalink's own parameter names rather than a second vocabulary invented for the API, and returns a strong ETag with a day of shared-cache life and a week of stale-while-revalidate.

A conditional request that still matches costs the caller a round trip and nothing else. POST is unchanged, and one shared reading of the inputs serves both so a link and the form behind it can never come to describe different configurations.

Decode speed, next to the figure the other calculator publishes

Memory is arithmetic and every serious tool lands close, so a memory error rate separates nobody. Decode speed does, and it is the number the nearest comparable calculator publishes as its own weakest: 15.1% median error and 1.22× geometric mean, against 2.6% and 1.04× here. The accuracy page now shows both, side by side, with the date they were read off that site's own page — and with the caveat stated at the same size as the claim, because two self-reported figures on two self-chosen corpora is not a controlled comparison and pretending otherwise would be the failure this page exists to avoid. What it does support is narrower and still worth saying: the cases here are scored leave-one-out, so every one is predicted by a calibration fitted without it, which is the only way a self-reported error rate means anything.

The optimisation that was measured and then not built

The 240-model catalogue reaches the browser as 122 KB of React payload, three quarters of it the per-model quantization lists, and encoding it as compact tuples cuts that to 22 KB. An 82% reduction is the kind of number that ends an argument, so it was measured on the wire before anything was written: after gzip the same change is 6.8 KB against 5.5 KB. Compressed size is the only size a reader ever pays for, and 1.3 KB does not buy a hand-rolled wire format and the decoder that comes with it, so it was not built. The measurements are published in the README along with the one finding that does matter and is not in this repository at all: the server compresses with gzip only, and the same homepage is 20 KB under brotli against 27 KB under gzip — a quarter off every document, for a setting at whatever terminates TLS.

2026-09-04

6 entries

The cross-check was accusing publishers of a silence that was ours

This site’s comparison page reported eighty-nine fields as “the publisher states nothing” across twelve models — the whole Qwen3.5 and Qwen3.6 generation, Kimi K3, MiniMax-M3 and both Inkling variants. The publishers stated all of it. A multimodal config.json nests the language model under text_config, and the comparison was reading only the top level, finding nothing, and recording the publisher as silent.

The catalogue had never had this bug: it stored the exact JSON Pointer it read each field from, and those pointers said /text_config/num_hidden_layers all along. Only the comparison ignored them. Both now resolve the publisher’s value through the recorded pointer, from one shared module, and the audit additionally fails if a published comparison cites a location the catalogue did not actually read — so a nested figure can never again be presented as a top-level one.

Comparable fields went from 211 to 280 and the unknowns from 89 to 20. The thirty-six places where the other calculator contradicts the publisher’s own file are unchanged, because those were never the doubtful part.

Being absent is its own kind of wrong, so it now breaks the build

Every audit here checked that a figure traced to its source. None checked whether the model somebody came to ask about was in the catalogue at all — and from a reader’s side, silence about a release is indistinguishable from being wrong about it. Coverage is now measured against two independent signals: the repositories a commercial inference market serves, and the top three hundred of Hugging Face’s own trending, download and like rankings.

Every eligible release from the current year that shows demand must be in the catalogue or recorded, with its reason, in a file anyone can read. A gap may exist; a gap nobody wrote down may not. The first run reported eighty-two, most of them requantizations and fine-tunes drowning the real ones.

Hugging Face records that relationship in a repository’s own tags, and all four kinds — quantized, fine-tune, merge, adapter — preserve geometry exactly, so the base already in the catalogue answers for them. Discounting those left eighteen genuine gaps, now closed: the catalogue holds 240 models, up from 200. The rule caught its first stale entry the same day, when Qwen3.8-Flash-Next started serving the config.json that had been answering 405 and the recorded reason stopped being true.

“A faster card would be faster” is true of almost everything, and useless

Decoding is bandwidth-bound at every speed, and there is nearly always a faster device in the catalogue, so the old rule fired on almost every configuration that fit. It told someone using a quarter of their RTX 4090 at 130 tokens a second to buy a workstation card for a 1.33× gain. There is a new verdict for what that actually is: nothing to fix.

It applies when the device has real memory headroom and already decodes faster than anyone reads, and the threshold — thirty tokens a second — is named, exported and printed in the verdict itself, because it is an editorial judgement rather than a measurement and anyone disagreeing should be able to see the number they disagree with. The faster device is still costed and still shown; it is offered as throughput somebody may want, not as the answer to a problem they do not have.

Power, published as the ceiling it actually is

Calculators print a wattage, a monthly bill and an annual CO₂ mass for every configuration, from defaults nobody sources. The only power figure a manufacturer publishes is the board’s own limit — TGP, TBP, or a module rating — and that is what the card may draw under a load designed to saturate it. Decoding is not such a load: generation waits on memory, so the arithmetic units idle and real draw sits below the limit by an amount no vendor publishes.

So this reports the limit, says plainly that it is a ceiling, and refuses to multiply it by an invented utilisation to manufacture a point estimate that would look more precise and mean less. The tariff and the grid’s carbon intensity are yours: electricity costs several times more in one country than another and grid intensity varies more than tenfold by region and by hour, so both stay blank until you supply them, exactly as the rental price already did.

The homepage now serves the answer instead of promising one

The verdict was fetched by the browser after the page had already loaded, so the most valuable thing on the site was absent from the HTML: a search engine indexed a calculator that appeared to hold no answers, and the largest paint waited on a round trip that was never going to say anything surprising. The opening configuration is now analysed while the page is built, and the browser skips that first request entirely. The empty panel that used to say “pick a model, a format and a device” — while a model, a format and a device were already picked — is gone. Three smaller things went with it: the evidence badge now says what share of the total was measured, so an answer whose weights and cache are both read from published files no longer reads as a guess because of an 0.80 GB runtime reserve; time-to-first-token shows the calibrated figure with the physical floor beside it, rather than a range that implied tenfold uncertainty about something that is not uncertain; and the four sliders stopped rendering in the browser’s default blue.

How long a conversation, answered per device

Every pair page answered at one reference context, which quietly hides the failure everyone eventually meets: the model loads, the chat grows, and the KV cache runs the device out of memory an hour later. Each pair now states the longest context it actually holds at the best format that fits, and says whether the wall is the card running out or the publisher’s own published maximum — only the first can be bought around. It is also the figure that makes each of these pages about its own pair rather than about its model, since the ceiling moves with the device.

2026-09-03

5 entries

The free tier gets a real ceiling, and a page that shows it

Two thousand calls a day was a gift, not a trial: enough to run a business on, which meant nobody would ever have a reason to pay. The free allowance is now two hundred a day — enough to read the documentation, wire up a client, test a dozen models and run something small, and deliberately not enough to serve a shop’s catalogue. Nothing was taken away except volume: every endpoint still answers on a free key, at full fidelity, with the same evidence tiers and the same refusals, because hiding a feature teaches a caller that the paid tier is a different product.

The ceiling is also now real rather than notional. It used to live in memory, which meant a restart refunded it and there was no answer to “what did I use this month”; the daily count is written to disk and is the authority, while the burst allowance stays in memory where a restart costs nobody anything. And it is visible: /usage shows exactly where a key stands, checking costs nothing against the allowance, and past eighty per cent of the day every response carries an X-RateLimit-Warning header.

A limit nobody can see reads as an outage.

Many configurations in one request, priced per configuration

A shop with four hundred cards on sale needs a table, and building it out of four hundred round trips is slow for them and wasteful here. The new batch endpoint takes up to ten configurations on the free plan and 250 on the largest, and each item costs one call — pricing a batch as a single call would make the quota meaningless the moment anybody noticed, and pricing it per HTTP request would mean a bill that tracked how somebody chose to wrap their work rather than the work. The allowance is checked for the whole batch before any of it runs, because a batch refused halfway would charge for analyses nobody received. One malformed item never fails the rest: every result carries its own status, so four hundred inputs return 399 answers and one named error.

Two thousand pages answering the question people actually type

“Can a 4090 run Qwen3?” is the query, and neither a model page nor a device page answers it — each knows half. There is now a page per pairing for the most downloaded models against the devices people own: which of the published formats fit, how fast each one goes, and, when none fit, the cheapest way out, including the option most sizing tools never mention, which is more of the same card. Two rules keep these from being the thin generated pages the technique is known for.

A pairing the engine cannot size is not published at all rather than published empty — a model selected purely on download count once put sixty 404s into the sitemap, and the selection is now the same predicate as the answer. And speed is quoted only for a format that actually fits: a rate printed beside “short by 60 GB” reads as a promise the configuration cannot keep.

The 8 GB tier exists, because that is what people own

The catalogue jumped from 12 GB straight to 16 GB and held nothing at all in the 8 GB range, which is the most common amount of video memory among exactly the people asking whether something will run. Every one of them was being told to pick a card they do not have. NVIDIA stopped publishing memory bandwidth for the 30 and 40 series on its web specification pages — capacity and interface width, nothing else — but it is still in the RTX Blackwell architecture whitepaper, whose appendix compares each 50-series card against its two predecessors.

The RTX 4070, 4070 Ti, 3080, 3070 Ti and 3070 are added from there, each figure checked twice: against the whitepaper’s own arithmetic, since interface width times data rate over eight reproduces every one of them exactly, and against three cards already catalogued from other sources that appear in the same tables and agree. The RTX 3060 12 GB and 4060 Ti 16 GB are still absent and stay absent: neither appears in that whitepaper, and taking their bandwidth from a third-party database would put an unsourced number underneath thousands of generated pages.

Written for somebody who does not already know the words

The front page opened with “what is actually slowing your LLM” and three panels about safetensors metadata, immutable revisions and block layouts. All true, all addressed to somebody who already understood the answer. It now asks the question a visitor actually has — will it run on your machine — and says in one sentence what they get: yes or no, how fast, and what to change.

Underneath, the three things an answer contains are explained without jargon, and the four terms that do most of the work on this site — quantization, context, KV cache, offload — are each given a plain sentence. Nothing about the rigour changed; what changed is that the page no longer requires the reader to have already won the argument.

2026-09-02

6 entries

A free API key, issued in the response that asked for it

The API was documented and unusable: it required a credential nobody could obtain. It now issues one immediately — an email address, a button, and the key is in the 201. No account, no confirmation link, no card, and every endpoint open on it including the cluster planner and the target inverter, because the paid tiers are meant to sell volume rather than capability.

Only the SHA-256 digest of a key is stored, so it is shown once and cannot be recovered by us or extracted from us. The address is not verified, and the interface says so rather than staging a “check your inbox” step behind a mail provider that does not exist. Budgets are a token bucket rather than a daily window: sixty calls back to back and two thousand a day, with the counters living in the serving process — which every response declares in X-RateLimit-Scope instead of leaving an integrator to discover it.

There is an OpenAPI 3.1 document at /v1/openapi.json, generated from the live catalogue so a client cannot be built against a model the engine no longer has.

The same figure, checked against the publisher’s own file

Every architecture number here is read from a pinned config.json. That claim is only worth something tested, so a script now fetches another public calculator’s specification tables and compares both sites against the publisher’s file, which is the arbiter — neither site is the referee. Across fifty models and 211 comparable fields the other source differs from the publisher on thirty-six; this catalogue differs on none.

The disagreements are not small: one widely used routed model is published there with sixty layers, ninety-six query heads, eight key-value heads and a hidden size of 4096 against the publisher’s forty-eight, thirty-two, four and 2048, which sizes its cache two and a half times too large. Two rules keep this a check rather than a complaint: a value the scraper cannot read is recorded as unread with the reason and never counted as a disagreement, and our own column is graded against the same file by a build-time audit, so a green score for this site has to be earned rather than asserted.

The answer comes before the caveat

The result panel opened with a verdict and, immediately under it, a sentence about an uncalibrated runtime reserve. Every word of that was true and it was the wrong order to read it in. A new panel now leads with the thing somebody came for — fits or does not fit, the share of the device it takes, tokens per second, time to first token and the memory left over — with the evidence badge beside it and the caveats below rather than in front. Nothing was removed and no claim was softened; only the reading order changed.

What people actually run, joined to what it takes to run it

A ranking of open-weight models by Hugging Face’s own thirty-day download count, cited with the URL and the date, joined to the smallest catalogued device that holds each one fully resident. The ordering is somebody else’s published measurement rather than this site’s traffic, which a reader could not falsify and which on a young site is mostly noise. Sizing is quoted at the format people download rather than at FP16, and the caveat says plainly what a download count is: a measure of attention on one host, not of quality. Every hardware page now answers the same question from the other side — the whole catalogue run against one device, largest that fits first.

Tailwind utilities were silently dead on every link

Tailwind v4 puts its own rules inside cascade layers, and an unlayered rule beats every layered one regardless of specificity. A bare a { color: inherit } in this stylesheet therefore defeated every text-colour utility on every anchor in the site, and .evidence-badge defeated Tailwind’s hidden — which is why a badge meant to disappear below 640 pixels kept rendering and pushed a 320-pixel screen sideways. Both looked correct in the source and neither worked in the browser.

The stylesheet now declares its layers: element resets in base, component classes in components, and the print sheet unlayered because it is the one place that must override everything. Separately, grid and flex children default to a minimum width of auto, which let one code sample or one wide table widen a whole track instead of scrolling inside its own box; that is now defaulted to zero through a zero-specificity rule any explicit style still beats. No page overflows horizontally at 320, 375, 768, 1280 or 1600 pixels.

The widget sizes itself, and there is a page for shops

The embed was a fixed 620 pixels, which is the reason most embeds look broken on somebody else’s product page — too short and the answer is cut off, too tall and there is a slab of empty background in the middle of a listing. The frame now measures its own content and reports it, and the loader applies it after checking both the origin and the source frame and clamping the value. It is measured on the body rather than the root element, because the root is stretched to the frame and can therefore only ever grow.

A mutation observer sits beside the resize observer, since browsers suspend rendering — and with it resize callbacks — in an offscreen frame, which is where a lazily-loaded embed spends most of its life. The new page for shops builds the script tag for a chosen device and shows it working in a live frame beside it rather than in a mock-up.

2026-09-01

9 entries

Nine files outside the band, settled by reading their headers

The weight scorecard fails the build when a published file falls outside the band the interface shows, because a band that misses a real file would mislead someone. Nine were missing, from three repositories, and widening every band until they fit would have degraded the answer for the 199 models whose publishers used the stock mixture. Reading the GGUF headers of anthracite-org/magnum-v4-72b settled it instead: its Q4_K_M puts ffn_down at Q8_0 on forty layers and Q5_0 on the other forty, where a stock Q4_K_M uses Q4_K with Q6_K on the raised subset, and its Q6_K puts ffn_down at Q8_0 on all eighty.

Those files are fatter than their names, which is also why the Q8_0 file of the same repository lands exactly — Q8_0 has no mixture to deviate from. The nine are now named with that evidence, the contract holds for everything else, and a tenth still fails the build.

State the outcome you need, get the configurations that reach it

The calculator answers what a device does. The new service-target endpoint answers the opposite question — what reaches a given speed, latency, device budget or cost per million tokens — which is the question someone actually has when they are buying. It returns three lists rather than one: what passes, what fails with every missed requirement named, and what could not be judged because a figure it needs is not published.

That third list is the point: a device with no published time to first token must not silently satisfy a latency target, and dropping it would make the passing list look more complete than the evidence is. Passes are ordered by device count, because the cheapest way to meet a requirement is the answer, not the fastest way.

Saved configurations, CSV and a print sheet

Configurations can be saved without an account, because asking someone to register to remember a dropdown would charge in attention for something that costs nothing to store locally. They live in one browser and the interface says so, since a saved thing that silently disappears is worse than one that never claimed to persist; each entry is the same permalink the calculator already produces, so a saved configuration and a shared link cannot drift apart. The CSV export carries the evidence level beside every figure and explains the levels in its header, because a spreadsheet outlives the tab it came from and a bare number in it would lose the only thing separating this from a guess. PDF is a print stylesheet rather than a dependency: ink on white, controls hidden, and every source URL expanded after its link.

Kimi-K3 stops being priced as if all 93 layers grew with context

Moonshot describes Kimi-K3's hybrid stack by listing which layers keep full attention inside linear_attn_config rather than by writing a layer_types array, so the engine had been treating all 93 layers as attention — a conservative bound, but a large one. That list carries the same information and is now expanded: 24 layers keep full attention and the other 69 hold a constant recurrent state that does not grow with context.

Cluster mode reaches the calculator, and it changes the verdict

The single rig / cluster switch is on the front page. Choosing more than one device divides the weights and the cache and leaves the whole runtime reserve on every card, which is what actually has to fit — Qwen3-32B at Q4_K_M is capacity bound on one RTX 4090 and bandwidth bound on two, and the verdict now says so instead of answering a question nobody asked. Decode scales with it, but only where the physics allows: tensor parallelism shards every layer so each device streams a quarter of the bytes per token and four cards run close to four times faster, while pipeline parallelism runs its stages in sequence and a single request sees no latency gain at all. The interconnect is charged rather than assumed free, so a cluster on 25 GbE is slower than the same cluster on NVLink.

Two figures that are yours, not ours

Usable memory per device can now be overridden, for a laptop part sharing memory with a display or a card the catalogue does not carry. And rental cost is published as dollars per million tokens — the only figure that lets a fast expensive device be compared with a slow cheap one — from an hourly price you supply. No price is kept here: the same card differs by more than threefold between providers, moves week to week, and has no publisher document to pin it to. The result says the price was yours and that an idle device still bills.

The engine now answers every catalogued configuration

The sweep used to declare 344 of 1,600 audited calculations unsupported. It now declares none, and not one figure was invented to get there. Every refusal turned out to be a published field the engine was not reading: gated delta net, Mamba-2 and short-convolution layers publish the dimensions their constant state is made of; LongCat, GLM-4.7-Flash and Kimi-K3 publish kv_lora_rank next to qk_rope_head_dim, which is what a compressed latent cache is, and were being refused only because the router matched on model_type strings instead of on geometry; Falcon publishes multi_query, which means one KV head and not seventy-one; and GLM writes its linear-attention head count inside a nested linear_attn_config object. Three models whose publisher states no context ceiling are now measured at the audit's own contexts, with the ceiling recorded as unpublished, instead of having the whole row refused over one missing field.

A recurrent layer's cache no longer grows with context

Hybrid models keep a fixed-size state on their linear-attention, Mamba and convolution layers, and modelling that as unsupported threw away the best news in the calculation: a 64-layer Qwen3.5 holds full attention on one layer in four and a constant state on the other three. The state is sized from the publisher's own dimensions, laid out as the reference implementation in transformers lays it out, and it is cited that way. It also does not shrink when the KV cache is quantized, which the result says.

Fine-tuning, cluster planning, prefix caching and speculative decoding

Four capabilities the calculator did not have. Fine-tuning splits into four terms and refuses to blend them: weights, gradients and optimizer state are exact arithmetic on a verified parameter count, using the ZeRO paper's 12 bytes per parameter and QLoRA's 4.127 bits, while activation memory is labelled as the soft term it is. LoRA adapters are counted from the real projection shapes rather than as a share of the model.

Cluster planning divides weights and cache by the device count and pointedly does not divide the runtime reserve, and prices the interconnect ceiling rather than a speedup nobody has measured. Speculative decoding asks the reader for an acceptance rate instead of inventing one.

2026-08-31

26 entries

Two quantizations were being added together

Discovery resolved a file's quantization by searching for a known name followed by a separator, and what follows the separator is usually the rest of the same name: Q6_K_L contains Q6_K followed by an underscore. A repository publishing both had the two summed into one artifact, recording Trinity-Large-Thinking Q6_K as 687 GB when it weighs 343, which put it 53% outside its own published band. The same defect applied to Q4_0_4_8 against Q4_0, despite the comment claiming to guard against exactly that.

The token is now read whole and then recognised, never searched for by prefix. Worst published-size error fell from 52.7% to 26.3%.

Nine files sit outside the derived band, and are listed

Three repositories ship mixtures that are not the standard one for the name on the file: a “Q4_K_M” of 6.34 effective bits per weight is not a stock Q4_K_M. In all three the parameter count reconciles and the Q8_0 file lands exactly, so the geometry is right and the mixture is not. Those nine are named in the test suite rather than absorbed by widening every band on the site, and they do not reach a reader: when a published file exists the engine serves its actual byte count. Band coverage of the derived tier is 95.7%.

Time to first token now covers 16 devices instead of 3

Prefill runs its matrix multiplications on the tensor or matrix units, so pricing its compute roof from the headline vector figure is wrong by the ratio between the two: a factor of two on an RTX 3090, four on an RTX 6000 Ada. Until now the estimate was published only where a vendor's own label said the figure was a vector one, which meant three AMD cards and nothing else. The catalogue now carries a dense matrix peak quoted from the vendor's specification table for sixteen devices, and the calibration is re-solved against the benchmarked card's own matrix peak of 191 TFLOPS, giving 9.58% utilisation where the vector solve gave 38.30%. Every published time to first token moved with it.

A sparsity figure is never halved into a dense one

NVIDIA quotes H200 at 1,979 FP16 teraFLOPS footnoted “With sparsity”, and the RTX PRO 6000 Blackwell at 4,000 AI TOPS footnoted “Theoretical FP4 TOPS using sparsity”. Those numbers are exactly twice the dense rate and describe a workload with half its weights pruned, which a dense transformer is not. Rather than divide by two and present the result as the vendor's, those devices carry no matrix peak and no time to first token, and each says so in its own caveat. The same applies to Apple silicon, which publishes no matrix throughput for the GPU at all.

Published headers corrected the low-bit mixtures

Q2_K was over-estimating every routed model by a 7.99% median with a systematic +7.28% bias. The transcribed rule raised the value and down projections on every layer; the published headers say Q2_K raises neither and holds both uniformly at Q3_K, while Q3_K_M raises them from Q4_K to Q5_K on the usual subset. Reading the real headers instead of inferring brought Q2_K to a 1.44% median with the bias gone.

A reconstruction now carries its own reconciliation gap

A model reproducing its published parameter count to 1.6% was being refused a reconstruction entirely and priced with block bounds instead, which made the most downloaded model on Hugging Face show a range twice as wide as it needed. Such a model is now reconstructed, and its own gap is added to the published range rather than ignored, so the range says how well that particular architecture is understood. Gemma 4 12B moved from a 5.1 to 12 GB range to 7.0 to 8.1 GB.

Two data bugs found by widening the corpus

A multimodal projector file was being read as a model's weights, recording a 397B model as weighing 0.9 GB. And one publisher's safetensors index reports a parameter count its own published bytes contradict by five orders of magnitude. Both are now refused: a projector is not weights, and a parameter count implying more than nine bytes per parameter is not usable.

The lower bound admits what a conversion drops

A published parameter count describes the whole checkpoint, but a GGUF holds only what llama.cpp runs: a multimodal vision tower and a multi-token-prediction head go into a separate file or nowhere. Ten published files sat below a bound that assumed otherwise. The bound now allows for it, measured at the largest observed shortfall.

Catalogue rebuilt: 160 models, 24 publishers

Discovery was ranking each publisher by downloads, which is dominated by older repositories and hid every recent release. It now pins the current flagships explicitly and covers publishers it had never queried. Gemma 4, Mistral Small 3.2, Magistral, Devstral, Xiaomi MiMo, Tencent Hy4, MiniMax M2.7, GLM-5.3 Flash, Nemotron 3.5, granite 4.2, Ling 3.0, Ornith 1.5 and ERNIE 4.5 are all in the snapshot for the first time.

Why Llama is not here

Every meta-llama and older google repository on Hugging Face sits behind a manual licence gate, and a gated repository does not serve its config.json: the request returns 401. Architecture is only accepted from the model's own publisher, so the Llama family cannot be catalogued without a licence acceptance this project does not hold. Gemma 4 is a different case and is included, because Google published it ungated.

Official GGUF coverage more than tripled

A discovery pass now looks for a GGUF repository owned by the same publisher as each catalogued model, so the exactly-known weight tier grew from 108 artifacts across 16 models to 335 across 54, spanning eleven publishers instead of one. Community requantizations are still never admitted: a file is only used when its own publisher released it.

The size model held up outside its training family

Until now the reconstruction had only ever been scored against Qwen files. Replayed against the enlarged corpus it lands at a 0.012% median across eleven publishers, which is the first evidence that it generalises rather than fitting one publisher's conventions.

Two config shapes that were being rejected

A published max_window_layers of zero means every layer is windowed; the validator was treating it as invalid and dropping the model. Hunyuan publishes moe_intermediate_size as one entry per layer, and MiMo publishes moe_layer_freq as a per-layer mask. Uniform per-layer arrays are now collapsed to the value they all carry, non-uniform ones are refused rather than averaged, and the per-layer mask is kept in its own field instead of being discarded.

Hardware: RTX 50 series, DGX Spark, M4 and M5 Pro

Added the RTX 5070 Ti, RTX 5070 and RTX 5060 Ti 16GB from NVIDIA's comparison specifications, DGX Spark from its product page, and the M4, M4 Pro and M5 Pro from Apple's own announcements, each field carrying the sentence that states it.

Two devices deliberately left out

AMD's product pages refused every automated request during this pass, so the RX 9070 non-XT, RX 9060 XT and RX 7800 XT wait for a specification we can quote from AMD directly. The Xiaomi AI Cube, announced on 24 August, publishes 80GB of unified memory and 1.22TB/s of near-memory computing bandwidth: that figure describes compute adjacent to memory on its accelerator, not the streaming bandwidth this engine's decode model consumes, and Xiaomi has published no specification page for it. Entering it would produce confident tokens-per-second figures with nothing behind them.

Every model now answers, in three tiers

The calculator previously offered only the sixteen models with a published GGUF. Weight files are now sized on a tier ladder: a published artifact where one exists, a reconstruction from the pinned architecture and exact ggml block layouts where none does, and block-geometry bounds where the architecture cannot be reconstructed. The tier travels with every number, and the reconstruction's measured error is published per format on the accuracy page.

Routed models are priced on the bytes a token reads

MoE decode was withheld entirely. It is now computed from the shared tensors plus the experts a token actually visits, with the expected number of distinct experts derived from independent routing at batch sizes above one. Capacity still requires the whole file resident, which is why a routed model can be fast and still not fit.

Time to first token is published as a range

Prefill utilisation is solved from one published benchmark rather than inherited from a competitor, so a point estimate would overstate what is known. The physical roofline floor and the utilisation-adjusted figure are shown together, and a device whose vendor publishes no compute peak gets a memory-roof lower bound instead of silence.

Offload and concurrency reach the interface

Layer offload, batch, concurrent users, runtime and declared system bandwidth are now inputs rather than a note saying they are unavailable. Offload is quantised to whole layers because that is what llama.cpp does. Free-plan ceilings clamp the result and name the capability that was clamped.

A speed verdict no longer inherits the runtime reserve

The uncalibrated runtime reserve is identical on both sides of a device comparison, so it cancels. Marking every bandwidth verdict a hypothesis because of a term that does not affect it was its own kind of dishonesty. The reserve still governs the evidence level of capacity verdicts, where it genuinely binds, and every result now names its weakest input.

Latest-model snapshot refreshed

Refreshed four model entries with pinned official configs: DeepSeek-V4-Pro-0813, GLM-5.3, LFM2-24B-A2B and LFM2-2.6B-Longevity. The catalog remains at 100 models; recency did not override source or coverage invariants.

DSA false precision blocked

DeepSeek V4 and GLM-MoE-DSA return unsupported for KV memory. Their official implementations maintain compressed, expanded or indexer state that the conventional KV-head formula does not represent.

Radeon AI PRO R9700 source added

Added AMD's official 32 GB workstation GPU specification: 640 GB/s bandwidth, 47.8 FP32 TFLOPS and 300 W. Its published llama.cpp results are used as a founded observation and, since this change, as the single anchor behind the prefill calibration.

RTX PRO 6000 Blackwell source added

Added the Workstation Edition from NVIDIA's official specification: 96 GB GDDR7 ECC, 1,792 GB/s, 125 single-precision TFLOPS and 600 W. Bus width remains blank because NVIDIA does not publish it on that source.

GGUF representation invariant

Corrected 20 artifacts where a monolithic GGUF and its equivalent multipart publication had been summed together. Added ingestion and runtime invariants that reject mixed representations.

Qwen3.8 MTP boundary

Recorded that the official config declares one MTP layer while the pinned standard Transformers model ignores MTP keys. That mismatch is why the model does not reconcile and is sized with bounds rather than a reconstruction.

2026-08-30

1 entry

Initial verified snapshot

Published 100 pinned model configs, 23 manufacturer-sourced hardware entries, exact official Qwen GGUF artifacts and the first narrow decode scorecard.

llmbottleneck
catalogue 2026-10-03models 327devices 135