The same engine, programmatically
For stores, configurators and local-LLM apps. Every response keeps the evidence tier and the reason for anything unsupported.
1 · Get a key
Keys are issued instantly
A key is issued at /keys: an email address, a button, and the key is in the response. Every endpoint is open on it, including the cluster planner and the target inverter.
| Free budget | 200 calls a day, 20 back to back |
|---|---|
| The free plan’s condition | Show Hardware compatibility data: LLM Bottleneck, linked to this site, beside results you display publicly. Internal tools need nothing; a paid plan removes the requirement. |
| Clients | TypeScript and Python one file each, no dependencies, generated from the live spec |
| Machine-readable spec | /v1/openapi.json OpenAPI 3.1, generated from the live catalogue |
| Check your usage | GET /v1/usage with your key, or the usage page. Costs nothing against the allowance. |
| Before you are refused | Past 80% of the day, every response carries an X-RateLimit-Warning header. |
| Billing | The free tier is not billed. Pro and Business remove the attribution requirement, unbrand the widget on your domains and raise the daily allowance; see the plans. |
2 · Your first call
Two commands. The key is in the first response and shown only once.
# 1. Take a key. It is in the response, and shown only once.
curl -s https://llmbottleneck.com/v1/keys \
-H "content-type: application/json" \
-d '{"email":"you@example.com","label":"first try"}'
# 2. Use it.
curl -s https://llmbottleneck.com/v1/analyze \
-H "Authorization: Bearer llmb_..." \
-H "content-type: application/json" \
-d '{"model":"qwen-qwen3-8b","quantization":"Q4_K_M",
"context":8192,"hardware":"nvidia-geforce-rtx-4090"}'Prefer a typed client? Generate one from /v1/openapi.json, which enumerates every valid model and device identifier from the same catalogue that answers.
3 · The response
POST /v1/analyze takes one configuration:
{
"model": "qwen-qwen3-8b",
"quantization": "Q4_K_M",
"context": 8192,
"hardware": "nvidia-geforce-rtx-4090"
}The response carries, for every field, how it was obtained:
{
"result": {
"weightSize": { "tier": "verified", "bytes": …, "low": …, "high": … },
"capacity": { "fits": true, "requiredBytes": …, "shortfallBytes": 0 },
"offload": { "active": false, "layersOnHost": 0 },
"speed": { "status": "ok", "tokensPerSecond": … },
"ttft": { "status": "ok", "floorMs": …, "calibratedMs": … },
"diagnosis": { "code": "BANDWIDTH_LIMITED", "recommendation": { … } }
}
}speed is the engine’s roofline estimate for the layout it sized. Read it together with capacity.fits and offload: a configuration that does not fit and does not offload does not run, whatever the roofline says.
4 · Errors, and what each one means
| Status | Code | What it means |
|---|---|---|
| 401 | NO_API_KEY | No Authorization header was sent. |
| 401 | UNAUTHORIZED | A key was sent and is not recognised. Keys are not recoverable, so a lost one is replaced rather than looked up. |
| 422 | unsupported | The configuration is understood but cannot be answered from published data. missingFields names exactly what is absent. An unknown model slug gets this response — never a silent substitution. |
| 429 | RATE_LIMITED | Budget exhausted. Retry-After says how many seconds to wait. |
| 413 | BODY_TOO_LARGE | The body exceeded the cap. It is refused while streaming, not after buffering. |
Endpoints
| /v1/analyze | One configuration: memory, offload split, decode, TTFT, verdict and recommendation. |
|---|---|
| /v1/models | The pinned model catalogue with its revisions and retrieval timestamps. |
| /v1/hardware | Manufacturer-sourced devices, each field carrying its own citation. |
| /v1/recommend | The smallest published device that clears a tokens-per-second target, or an explicit refusal when none does. |
| /v1/finetune | Training memory for a full, LoRA or QLoRA run, split into the three terms that are arithmetic and the one that is not. |
| /v1/cluster | A model split across devices: what each one holds, and the ceiling the interconnect imposes. |
| /v1/batch | Up to 10 configurations per request on the free tier. Each item costs one call, checked for the whole batch before any of it runs; one bad item never fails the rest. |
| /v1/usage | What this key has spent today and over sixty days. Never metered. |
| /v1/changes | Models and devices added or changed since a date. Poll it weekly and an integration never goes stale; one call whatever the window. |
| /v1/resolve | Which catalogued model answers for any Hugging Face repository: a fine-tune, a merge or a GGUF requantization, with the evidence for it and the memory each published file needs. |
| /v1/widget/domains | The domains your widget renders unbranded on (Pro: one, Business: ten, with your own button). Never metered. |
| /v1/target | The inverse question: state the speed, latency, device count or cost you need, and get every configuration that meets it, fails it, or could not be judged. |
Embeddable widget
One script tag renders the free diagnosis inside a sandboxed iframe, so a storefront never handles a credential and never inherits our CSS. It needs no key. An attribute the frame cannot use is named in the frame rather than replaced.
<script src="https://llmbottleneck.com/widget.js" data-model="qwen-qwen3-8b" data-quantization="Q4_K_M" data-hardware="nvidia-geforce-rtx-4090" data-context="8192"></script>
On the free plan a one-line credit appears under the frame, linked to this site, and the frame links back to the full calculator. On Pro and Business, register your domains in Usage & keys and the widget there renders without our branding — and on Business with your own button, such as a store search or an affiliate link filled in with the model and device the reader has chosen.