LLMBOTTLENECK.COM
API v1

The same engine, programmatically

For stores, configurators and local-LLM apps. Every response keeps the evidence tier and the reason for anything unsupported.

1 · Get a key

Keys are issued instantly

A key is issued at /keys: an email address, a button, and the key is in the response. Every endpoint is open on it, including the cluster planner and the target inverter.

Get an API key →

Free budget200 calls a day, 20 back to back
The free plan’s conditionShow Hardware compatibility data: LLM Bottleneck, linked to this site, beside results you display publicly. Internal tools need nothing; a paid plan removes the requirement.
ClientsTypeScript and Python one file each, no dependencies, generated from the live spec
Machine-readable spec/v1/openapi.json OpenAPI 3.1, generated from the live catalogue
Check your usageGET /v1/usage with your key, or the usage page. Costs nothing against the allowance.
Before you are refusedPast 80% of the day, every response carries an X-RateLimit-Warning header.
BillingThe free tier is not billed. Pro and Business remove the attribution requirement, unbrand the widget on your domains and raise the daily allowance; see the plans.

2 · Your first call

Two commands. The key is in the first response and shown only once.

bashExample: request a key, then call /v1/analyze
# 1. Take a key. It is in the response, and shown only once.
curl -s https://llmbottleneck.com/v1/keys \
  -H "content-type: application/json" \
  -d '{"email":"you@example.com","label":"first try"}'

# 2. Use it.
curl -s https://llmbottleneck.com/v1/analyze \
  -H "Authorization: Bearer llmb_..." \
  -H "content-type: application/json" \
  -d '{"model":"qwen-qwen3-8b","quantization":"Q4_K_M",
       "context":8192,"hardware":"nvidia-geforce-rtx-4090"}'

Prefer a typed client? Generate one from /v1/openapi.json, which enumerates every valid model and device identifier from the same catalogue that answers.

3 · The response

POST /v1/analyze takes one configuration:

jsonRequest body
{
  "model": "qwen-qwen3-8b",
  "quantization": "Q4_K_M",
  "context": 8192,
  "hardware": "nvidia-geforce-rtx-4090"
}

The response carries, for every field, how it was obtained:

jsonResponse shape
{
  "result": {
    "weightSize": { "tier": "verified", "bytes": …, "low": …, "high": … },
    "capacity":   { "fits": true, "requiredBytes": …, "shortfallBytes": 0 },
    "offload":    { "active": false, "layersOnHost": 0 },
    "speed":      { "status": "ok", "tokensPerSecond": … },
    "ttft":       { "status": "ok", "floorMs": …, "calibratedMs": … },
    "diagnosis":  { "code": "BANDWIDTH_LIMITED", "recommendation": { … } }
  }
}

speed is the engine’s roofline estimate for the layout it sized. Read it together with capacity.fits and offload: a configuration that does not fit and does not offload does not run, whatever the roofline says.

4 · Errors, and what each one means

StatusCodeWhat it means
401NO_API_KEYNo Authorization header was sent.
401UNAUTHORIZEDA key was sent and is not recognised. Keys are not recoverable, so a lost one is replaced rather than looked up.
422unsupportedThe configuration is understood but cannot be answered from published data. missingFields names exactly what is absent. An unknown model slug gets this response — never a silent substitution.
429RATE_LIMITEDBudget exhausted. Retry-After says how many seconds to wait.
413BODY_TOO_LARGEThe body exceeded the cap. It is refused while streaming, not after buffering.

Endpoints

/v1/analyzeOne configuration: memory, offload split, decode, TTFT, verdict and recommendation.
/v1/modelsThe pinned model catalogue with its revisions and retrieval timestamps.
/v1/hardwareManufacturer-sourced devices, each field carrying its own citation.
/v1/recommendThe smallest published device that clears a tokens-per-second target, or an explicit refusal when none does.
/v1/finetuneTraining memory for a full, LoRA or QLoRA run, split into the three terms that are arithmetic and the one that is not.
/v1/clusterA model split across devices: what each one holds, and the ceiling the interconnect imposes.
/v1/batchUp to 10 configurations per request on the free tier. Each item costs one call, checked for the whole batch before any of it runs; one bad item never fails the rest.
/v1/usageWhat this key has spent today and over sixty days. Never metered.
/v1/changesModels and devices added or changed since a date. Poll it weekly and an integration never goes stale; one call whatever the window.
/v1/resolveWhich catalogued model answers for any Hugging Face repository: a fine-tune, a merge or a GGUF requantization, with the evidence for it and the memory each published file needs.
/v1/widget/domainsThe domains your widget renders unbranded on (Pro: one, Business: ten, with your own button). Never metered.
/v1/targetThe inverse question: state the speed, latency, device count or cost you need, and get every configuration that meets it, fails it, or could not be judged.

Embeddable widget

One script tag renders the free diagnosis inside a sandboxed iframe, so a storefront never handles a credential and never inherits our CSS. It needs no key. An attribute the frame cannot use is named in the frame rather than replaced.

htmlWidget script tag
<script src="https://llmbottleneck.com/widget.js"
  data-model="qwen-qwen3-8b"
  data-quantization="Q4_K_M"
  data-hardware="nvidia-geforce-rtx-4090"
  data-context="8192"></script>

On the free plan a one-line credit appears under the frame, linked to this site, and the frame links back to the full calculator. On Pro and Business, register your domains in Usage & keys and the widget there renders without our branding — and on Business with your own button, such as a store search or an affiliate link filled in with the model and device the reader has chosen.

Build the tag for your own product →

llmbottleneck
catalogue 2026-10-03models 327devices 135