Ask your assistant, not a guess
Assistants answer “will this model run on my GPU?” from memory, and memory is months old. Connect this site and they answer from the catalogue — with the memory each format needs, a speed estimate, and a link to the page that shows the sources.
MCP server
One URL, no install. It works without a key for 20 calls a day; add a free key as Authorization: Bearer for 200 a day, or a paid plan for more.
https://llmbottleneck.com/mcp
Claude
Settings → Connectors → Add custom connector → paste the URL. In Claude Code:
claude mcp add --transport http llmbottleneck https://llmbottleneck.com/mcp
ChatGPT
Settings → Apps & Connectors → Advanced → Developer mode → Create → paste the URL, no authentication.
Cursor, VS Code, Windsurf and other clients
{
"mcpServers": {
"llmbottleneck": { "url": "https://llmbottleneck.com/mcp" }
}
}A client that only runs local commands can use the stdio bridge — the same file as the terminal tool below, at the path you saved it:
{
"mcpServers": {
"llmbottleneck": { "command": "node", "args": ["/path/to/llmbottleneck.mjs", "mcp"] }
}
}Tools
can_i_run— Can this device run this model?- Whether an open-weight LLM fits in a GPU's or Mac's memory, at every quantization or one you name, with the memory needed and an estimated decode speed. Sizes come from published files and manufacturer specifications.
what_runs_on— What runs on this device?- The open-weight models a GPU or Mac holds entirely in its memory, with the format, memory needed and estimated speed.
hardware_for_model— What hardware runs this model?- The single GPUs and Macs that hold a model at a quantization (default Q4_K_M), smallest memory first, optionally only those reaching a decode speed.
cheapest_way_to_run— Cheapest way to run this model- The cheapest rented GPU machine that holds the model (Vast.ai, RunPod prices) and the per-token API price, with a link to compare owning, renting and per-token for your own usage.
search_catalogue— Search the catalogue- Find the exact ids of catalogued models and devices from a partial name.
Every tool is read-only and answers from the pinned catalogue. Each answer ends with the page it came from, so what the assistant tells you can be checked.
In a terminal
Node 18 or newer. One file to download, no dependencies and nothing to install:
curl -fsSL https://llmbottleneck.com/v1/clients/llmbottleneck.mjs -o llmbottleneck.mjs node llmbottleneck.mjs # detect your GPU, list what it runs node llmbottleneck.mjs can-i-run "Qwen3-32B" # every format on the detected GPU node llmbottleneck.mjs hardware "Llama 3.3 70B" --min-tps 20 node llmbottleneck.mjs cheapest "Qwen3-32B"
Help calibrate the speeds. With llama.cpp installed, node llmbottleneck.mjs bench --model file.gguf measures your machine and shows the estimate beside it; --submit sends the numbers for review. Nothing is published without review, and only the GPU name, the file name and size and the benchmark’s figures are sent.
In code
The JSON API with typed TypeScript and Python clients, batch requests and a change feed is documented on the API page. For a product page, the widget is one script tag.