LLMBOTTLENECK.COM
Read from the providers' own APIs

Rent a GPU for your LLM

What each card rents for right now, what it runs, and what to rent for the models people download most — ranked by the providers' own live prices, never by who pays us.

Every rentable machine that holds it, ranked by live price, with speed, cost per token and the per-token API alternative.

Most downloaded models

For each, the rented machine this site recommends — the lowest cost per token among machines that answer at 20+ tok/s — and the cheapest per-token price where an API serves the model.

Speeds are this site’s estimates, not benchmarks of the rented machine. They are scored against benchmarks of one dense 7B model; for mixture-of-experts models and for several cards together there is no corpus to score them against yet. What is and is not measured

Prices · 18:30 UTC, 22 Sept

Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank

Every rentable card

One GPU, by the hour, right now

The cheapest in-stock price for a single card on each provider, next to what the card does with three reference models at Q4_K_M. Memory decides what fits; bandwidth decides how fast it answers.

Prices · 18:30 UTC, 22 Sept
GPUPrice nowOther offers$ / GB·hSpeed at Q4_K_MRent
RTX 306012 GB · 360 GB/s$0.036Vast.ai—$0.0038B~41 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 30708 GB · 448 GB/s$0.082Vast.aiRunPod Community $0.13$0.0108B~51 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 308010 GB · 760 GB/s$0.082Vast.ai—$0.0088B~87 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 4060 Ti 16GB16 GB · 288 GB/s$0.090Vast.ai—$0.0068B~33 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 407012 GB · 504 GB/s$0.096Vast.ai—$0.0088B~58 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 309024 GB · 936 GB/s$0.12Vast.aiRunPod Secure $0.50$0.0058B~108 tok/s32B~29 tok/s70Btoo bigRent on Vast.ai
RTX 5060 Ti 16GB16 GB · 448 GB/s$0.12Vast.ai—$0.0088B~51 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 3080 Ti12 GB · 912 GB/s$0.14Vast.ai—$0.0118B~105 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 507012 GB · 672 GB/s$0.18Vast.ai—$0.0158B~77 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 3090 Ti24 GB · 1008 GB/s$0.19Vast.aiRunPod Community $0.27$0.0088B~116 tok/s32B~31 tok/s70Btoo bigRent on Vast.ai
RTX 5070 Ti16 GB · 896 GB/s$0.20Vast.ai—$0.0138B~103 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 4080 SUPER16 GB · 736 GB/s$0.23Vast.ai—$0.0148B~85 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 508016 GB · 960 GB/s$0.25Vast.aiRunPod Community $0.39$0.0168B~110 tok/s32Btoo big70Btoo bigRent on Vast.ai
L424 GB · 300 GB/s$0.26Vast.aiRunPod Secure $0.49$0.0118B~34 tok/s32B~9.4 tok/s70Btoo bigRent on Vast.ai
RTX 6000 Ada48 GB · 960 GB/s$0.26Vast.aiRunPod Secure $0.84$0.0058B~110 tok/s32B~30 tok/s70B~14 tok/sRent on Vast.ai
RTX 408016 GB · 716.8 GB/s$0.27Vast.ai—$0.0178B~82 tok/s32Btoo big70Btoo bigRent on Vast.ai
RTX 509032 GB · 1792 GB/s$0.27Vast.aiRunPod Secure $0.99$0.0088B~207 tok/s32B~56 tok/s70Btoo bigRent on Vast.ai
RTX A600048 GB · 768 GB/s$0.28Vast.aiRunPod Secure $0.53$0.0068B~88 tok/s32B~24 tok/s70B~11 tok/sRent on Vast.ai
RTX 409024 GB · 1008 GB/s$0.34RunPod CommunityVast.ai $0.36RunPod Secure $0.74$0.0148B~116 tok/s32B~31 tok/s70Btoo bigRent on RunPod
RTX PRO 500048 GB · 1344 GB/s$0.51Vast.aiRunPod Community $0.82$0.0118B~155 tok/s32B~42 tok/s70B~20 tok/sRent on Vast.ai
L40S48 GB · 864 GB/s$0.69Vast.aiRunPod Community $0.79$0.0148B~99 tok/s32B~27 tok/s70B~13 tok/sRent on Vast.ai
RTX PRO 600096 GB · 1792 GB/s$1.14Vast.aiRunPod Community $1.69RunPod Secure $2.19$0.0128B~207 tok/s32B~56 tok/s70B~27 tok/sRent on Vast.ai
A100 80GB PCIe80 GB · 1935 GB/s$1.19RunPod CommunityRunPod Secure $1.59$0.0158B~223 tok/s32B~61 tok/s70B~29 tok/sRent on RunPod
A100 80GB SXM80 GB · 2039 GB/s$1.39RunPod CommunityRunPod Secure $1.59$0.0178B~235 tok/s32B~64 tok/s70B~31 tok/sRent on RunPod
H100 SXM80 GB · 3350 GB/s$1.74Vast.aiRunPod Secure $3.49$0.0228B~387 tok/s32B~106 tok/s70B~51 tok/sRent on Vast.ai
H100 PCIe80 GB · 2000 GB/s$2.40Vast.ai—$0.0308B~231 tok/s32B~63 tok/s70B~30 tok/sRent on Vast.ai
H200141 GB · 4800 GB/s$3.59RunPod CommunityRunPod Secure $4.59$0.0258B~554 tok/s32B~151 tok/s70B~73 tok/sRent on RunPod
B300270 GB · 7700 GB/s$7.89RunPod Secure—$0.0298B~889 tok/s32B~243 tok/s70B~117 tok/sRent on RunPod
RTX 4070 Ti12 GB · 504 GB/sout of stock——8B~58 tok/s32Btoo big70Btoo big
RTX PRO 5000 72GB72 GB · 1344 GB/sout of stock——8B~155 tok/s32B~42 tok/s70B~20 tok/s
H200 NVL141 GB · 4800 GB/sout of stock——8B~554 tok/s32B~151 tok/s70B~73 tok/s
B200180 GB · 7700 GB/sout of stock——8B~889 tok/s32B~243 tok/s70B~117 tok/s
MI300X192 GB · 5300 GB/sout of stock——8B~701 tok/s32B~192 tok/s70B~92 tok/s

One card, on-demand. Speeds are this site’s single-stream decode estimates at 8,192 tokens of context with the whole model in memory; “does not fit” means one card does not hold it — several may, and each model’s page prices those too. Vast hosts below 98% measured reliability are left out.

Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank

Choosing a card to rent

Memory first. A model runs at full speed only when all of it — weights, the KV cache for your context, and the runtime’s own reserve — fits in the card’s memory. Every table on this site answers that before it quotes a price, so nothing listed will spill to system RAM. At Q4_K_M and 8,192 tokens, Qwen3-8B needs 7 GB, Qwen3-32B 23 GB and Llama 3.3 70B 46 GB — two 24 GB cards or one 48 GB card for the last.

Then bandwidth. Generating each token reads every active weight once, so decode speed follows memory bandwidth far more than it follows compute. That is why an RTX 3090 and an A6000 decode at similar speeds, and why an H100 is several times faster than either on the same model.

Then the tier. Vast.ai is an open marketplace: independent hosts set their prices and Vast measures each host’s reliability, which is shown next to every price here (hosts below 98% are left out). RunPod Community Cloud is vetted third-party hosts; RunPod Secure Cloud is Tier III/IV data-centre capacity at a higher price. All three bill by the second and cost nothing once stopped, though a stopped machine’s disk is usually still billed.

Renting a GPU, or paying per token?

For one person chatting with a popular model, a per-token API is almost always cheaper: the provider batches your requests with everybody else’s, so you pay for tokens, not for a card that sits idle between your messages. Every model page shows the ratio for that model. A rented GPU wins when you keep it busy — batch jobs, many users, long agent runs — when the data must stay on a machine you control, when you need a fine-tune or a format no API serves, or when you want to test before you buy the card.

Starting a model on the rented machine

On either provider, create the machine from a container image and give it a start command. Each model page gives both for the format it priced — for llama.cpp, the image ghcr.io/ggml-org/llama.cpp:server-cuda with -hf <repository>:<format> -c 8192 -ngl 999 --host 0.0.0.0 --port 8080, which downloads the exact GGUF file the page sized and serves an OpenAI-compatible API on port 8080.

Not sure renting beats buying, or paying by the token? Compare a year of any model three ways →

How we rank, and who pays us

The order is the price; the recommendation is the arithmetic. Every cloud table sorts by the hourly price the provider quoted, among machines that hold the whole model, and you can re-sort by cost per token or by speed. Where one machine is marked recommended, the rule is written next to it and is the same everywhere: the lowest cost per million tokens among machines that answer at 20 tokens a second or more. Neither the sort nor the rule reads which provider pays us. A provider that pays nothing ranks above one that pays whenever the figures say so — OpenRouter, which pays us nothing, is listed first whenever its per-token price is lowest.

Some links earn a referral. The buttons to the providers below pay this site a share of what a new customer spends. It costs the customer nothing — the price is the same, and a new RunPod account gets a one-time $5 credit on its first $10 — and it is how the calculator stays free. Clicks are counted by provider and by page, never by person.

Store links, too. Where a page names a card the calculation chose — a larger one that holds the model, or the card a device page is about — it may link an Amazon or eBay search for it, and a purchase through one may pay this site a small commission. The link is attached after the card is chosen; no store pays for a position, and no store price is shown because a search page has none this site can vouch for.

ProviderWhat it pays usFor how longWhere its prices come from
Vast.ai3% of what a referred account spends. Termsfor the life of the accounthttps://console.vast.ai/api/v0/bundles/
RunPod3% of Pod spend and 5% of Serverless spend, in credits; 12% in cash once 25 referred accounts have paid. Termsthe first six months of each referred accounthttps://api.runpod.io/graphql
Novita AI10% of undiscounted credit purchases. Termsthe first 180 days, 60-day cookiehttps://api.novita.ai/openai/v1/models
OpenRouterNothing—https://openrouter.ai/api/v1/models

Considered and left out, each for a reason that is checked again by 2026-12-22:

  • DigitalOcean (10% for 12 months, cash (Awin)): Its GPU prices are rarely the cheapest for a card (H100 at $4.41/h on 2026-09-22), so an honest ranking would almost never send anyone there. Not integrated.
  • Massed Compute (up to 10% on approval, 20% discount code for readers): No public price API, and the approved terms (duration, payout) are unpublished. Add once both are in writing.
  • Vultr ($35 or $100 per active customer): Almost its whole GPU catalogue was out of stock on 2026-09-22 and it has no public price feed.
  • TensorDock, Lium, Verda, Jarvislabs, Hyperstack (none found): No referral programme and no unauthenticated price API to read; an unsourced price is not shown.
  • Thunder Compute, Hyperbolic, Spheron, Clore.ai (credits only or token amounts): Nothing a reader gains from, and nothing withdrawable.
  • Hetzner, Cudo Compute (closed): Hetzner ended its referral programme in June 2026; Cudo's affiliate page returns 404.

The programme terms were read from each provider on 2026-09-22. Who runs this site.

llmbottleneck
catalogue 2026-10-03models 327devices 135