$0.12per hour
- Speed
- ~239 tok/s
- By the token
- $0.30/M
Vast.ai · Marketplace · 99.5% reliable
All options →LLM//BOTTLENECK
What each card rents for right now, what it runs, and what to rent for the models people download most — ranked by the providers' own live prices, never by who pays us.
For each, the rented machine this site recommends — the lowest cost per token among machines that answer at 20+ tok/s — and the cheapest per-token price where an API serves the model.
Speeds are this site’s estimates, not benchmarks of the rented machine. They are scored against benchmarks of one dense 7B model; for mixture-of-experts models and for several cards together there is no corpus to score them against yet. What is and is not measured
$0.12per hour
Vast.ai · Marketplace · 99.5% reliable
All options →$0.12per hour
Vast.ai · Marketplace · 99.5% reliable
All options →$0.12per hour
Vast.ai · Marketplace · 99.5% reliable
All options →$0.12per hour
Vast.ai · Marketplace · 99.5% reliable
All options →$3.73per hour
Vast.ai · Marketplace · 99.9% reliable
All options →$0.12per hour
Vast.ai · Marketplace · 99.5% reliable
All options →$1.87per hour
Vast.ai · Marketplace · 99.2% reliable
All options →$1.07per hour
Vast.ai · Marketplace · 99.6% reliable
All options →Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
The cheapest in-stock price for a single card on each provider, next to what the card does with three reference models at Q4_K_M. Memory decides what fits; bandwidth decides how fast it answers.
| GPU | Price now | Other offers | $ / GB·h | Speed at Q4_K_M | Rent |
|---|---|---|---|---|---|
| RTX 306012 GB · 360 GB/s | $0.036Vast.ai | — | $0.003 | 8B~41 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 30708 GB · 448 GB/s | $0.082Vast.ai | RunPod Community $0.13 | $0.010 | 8B~51 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 308010 GB · 760 GB/s | $0.082Vast.ai | — | $0.008 | 8B~87 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 4060 Ti 16GB16 GB · 288 GB/s | $0.090Vast.ai | — | $0.006 | 8B~33 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 407012 GB · 504 GB/s | $0.096Vast.ai | — | $0.008 | 8B~58 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 309024 GB · 936 GB/s | $0.12Vast.ai | RunPod Secure $0.50 | $0.005 | 8B~108 tok/s32B~29 tok/s70Btoo big | Rent on Vast.ai |
| RTX 5060 Ti 16GB16 GB · 448 GB/s | $0.12Vast.ai | — | $0.008 | 8B~51 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 3080 Ti12 GB · 912 GB/s | $0.14Vast.ai | — | $0.011 | 8B~105 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 507012 GB · 672 GB/s | $0.18Vast.ai | — | $0.015 | 8B~77 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 3090 Ti24 GB · 1008 GB/s | $0.19Vast.ai | RunPod Community $0.27 | $0.008 | 8B~116 tok/s32B~31 tok/s70Btoo big | Rent on Vast.ai |
| RTX 5070 Ti16 GB · 896 GB/s | $0.20Vast.ai | — | $0.013 | 8B~103 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 4080 SUPER16 GB · 736 GB/s | $0.23Vast.ai | — | $0.014 | 8B~85 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 508016 GB · 960 GB/s | $0.25Vast.ai | RunPod Community $0.39 | $0.016 | 8B~110 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| L424 GB · 300 GB/s | $0.26Vast.ai | RunPod Secure $0.49 | $0.011 | 8B~34 tok/s32B~9.4 tok/s70Btoo big | Rent on Vast.ai |
| RTX 6000 Ada48 GB · 960 GB/s | $0.26Vast.ai | RunPod Secure $0.84 | $0.005 | 8B~110 tok/s32B~30 tok/s70B~14 tok/s | Rent on Vast.ai |
| RTX 408016 GB · 716.8 GB/s | $0.27Vast.ai | — | $0.017 | 8B~82 tok/s32Btoo big70Btoo big | Rent on Vast.ai |
| RTX 509032 GB · 1792 GB/s | $0.27Vast.ai | RunPod Secure $0.99 | $0.008 | 8B~207 tok/s32B~56 tok/s70Btoo big | Rent on Vast.ai |
| RTX A600048 GB · 768 GB/s | $0.28Vast.ai | RunPod Secure $0.53 | $0.006 | 8B~88 tok/s32B~24 tok/s70B~11 tok/s | Rent on Vast.ai |
| RTX 409024 GB · 1008 GB/s | $0.34RunPod Community | Vast.ai $0.36RunPod Secure $0.74 | $0.014 | 8B~116 tok/s32B~31 tok/s70Btoo big | Rent on RunPod |
| RTX PRO 500048 GB · 1344 GB/s | $0.51Vast.ai | RunPod Community $0.82 | $0.011 | 8B~155 tok/s32B~42 tok/s70B~20 tok/s | Rent on Vast.ai |
| L40S48 GB · 864 GB/s | $0.69Vast.ai | RunPod Community $0.79 | $0.014 | 8B~99 tok/s32B~27 tok/s70B~13 tok/s | Rent on Vast.ai |
| RTX PRO 600096 GB · 1792 GB/s | $1.14Vast.ai | RunPod Community $1.69RunPod Secure $2.19 | $0.012 | 8B~207 tok/s32B~56 tok/s70B~27 tok/s | Rent on Vast.ai |
| A100 80GB PCIe80 GB · 1935 GB/s | $1.19RunPod Community | RunPod Secure $1.59 | $0.015 | 8B~223 tok/s32B~61 tok/s70B~29 tok/s | Rent on RunPod |
| A100 80GB SXM80 GB · 2039 GB/s | $1.39RunPod Community | RunPod Secure $1.59 | $0.017 | 8B~235 tok/s32B~64 tok/s70B~31 tok/s | Rent on RunPod |
| H100 SXM80 GB · 3350 GB/s | $1.74Vast.ai | RunPod Secure $3.49 | $0.022 | 8B~387 tok/s32B~106 tok/s70B~51 tok/s | Rent on Vast.ai |
| H100 PCIe80 GB · 2000 GB/s | $2.40Vast.ai | — | $0.030 | 8B~231 tok/s32B~63 tok/s70B~30 tok/s | Rent on Vast.ai |
| H200141 GB · 4800 GB/s | $3.59RunPod Community | RunPod Secure $4.59 | $0.025 | 8B~554 tok/s32B~151 tok/s70B~73 tok/s | Rent on RunPod |
| B300270 GB · 7700 GB/s | $7.89RunPod Secure | — | $0.029 | 8B~889 tok/s32B~243 tok/s70B~117 tok/s | Rent on RunPod |
| RTX 4070 Ti12 GB · 504 GB/s | out of stock | — | — | 8B~58 tok/s32Btoo big70Btoo big | |
| RTX PRO 5000 72GB72 GB · 1344 GB/s | out of stock | — | — | 8B~155 tok/s32B~42 tok/s70B~20 tok/s | |
| H200 NVL141 GB · 4800 GB/s | out of stock | — | — | 8B~554 tok/s32B~151 tok/s70B~73 tok/s | |
| B200180 GB · 7700 GB/s | out of stock | — | — | 8B~889 tok/s32B~243 tok/s70B~117 tok/s | |
| MI300X192 GB · 5300 GB/s | out of stock | — | — | 8B~701 tok/s32B~192 tok/s70B~92 tok/s |
One card, on-demand. Speeds are this site’s single-stream decode estimates at 8,192 tokens of context with the whole model in memory; “does not fit” means one card does not hold it — several may, and each model’s page prices those too. Vast hosts below 98% measured reliability are left out.
Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
Memory first. A model runs at full speed only when all of it — weights, the KV cache for your context, and the runtime’s own reserve — fits in the card’s memory. Every table on this site answers that before it quotes a price, so nothing listed will spill to system RAM. At Q4_K_M and 8,192 tokens, Qwen3-8B needs 7 GB, Qwen3-32B 23 GB and Llama 3.3 70B 46 GB — two 24 GB cards or one 48 GB card for the last.
Then bandwidth. Generating each token reads every active weight once, so decode speed follows memory bandwidth far more than it follows compute. That is why an RTX 3090 and an A6000 decode at similar speeds, and why an H100 is several times faster than either on the same model.
Then the tier. Vast.ai is an open marketplace: independent hosts set their prices and Vast measures each host’s reliability, which is shown next to every price here (hosts below 98% are left out). RunPod Community Cloud is vetted third-party hosts; RunPod Secure Cloud is Tier III/IV data-centre capacity at a higher price. All three bill by the second and cost nothing once stopped, though a stopped machine’s disk is usually still billed.
For one person chatting with a popular model, a per-token API is almost always cheaper: the provider batches your requests with everybody else’s, so you pay for tokens, not for a card that sits idle between your messages. Every model page shows the ratio for that model. A rented GPU wins when you keep it busy — batch jobs, many users, long agent runs — when the data must stay on a machine you control, when you need a fine-tune or a format no API serves, or when you want to test before you buy the card.
On either provider, create the machine from a container image and give it a start command. Each model page gives both for the format it priced — for llama.cpp, the image ghcr.io/ggml-org/llama.cpp:server-cuda with -hf <repository>:<format> -c 8192 -ngl 999 --host 0.0.0.0 --port 8080, which downloads the exact GGUF file the page sized and serves an OpenAI-compatible API on port 8080.
Not sure renting beats buying, or paying by the token? Compare a year of any model three ways →
The order is the price; the recommendation is the arithmetic. Every cloud table sorts by the hourly price the provider quoted, among machines that hold the whole model, and you can re-sort by cost per token or by speed. Where one machine is marked recommended, the rule is written next to it and is the same everywhere: the lowest cost per million tokens among machines that answer at 20 tokens a second or more. Neither the sort nor the rule reads which provider pays us. A provider that pays nothing ranks above one that pays whenever the figures say so — OpenRouter, which pays us nothing, is listed first whenever its per-token price is lowest.
Some links earn a referral. The buttons to the providers below pay this site a share of what a new customer spends. It costs the customer nothing — the price is the same, and a new RunPod account gets a one-time $5 credit on its first $10 — and it is how the calculator stays free. Clicks are counted by provider and by page, never by person.
Store links, too. Where a page names a card the calculation chose — a larger one that holds the model, or the card a device page is about — it may link an Amazon or eBay search for it, and a purchase through one may pay this site a small commission. The link is attached after the card is chosen; no store pays for a position, and no store price is shown because a search page has none this site can vouch for.
| Provider | What it pays us | For how long | Where its prices come from |
|---|---|---|---|
| Vast.ai | 3% of what a referred account spends. Terms | for the life of the account | https://console.vast.ai/api/v0/bundles/ |
| RunPod | 3% of Pod spend and 5% of Serverless spend, in credits; 12% in cash once 25 referred accounts have paid. Terms | the first six months of each referred account | https://api.runpod.io/graphql |
| Novita AI | 10% of undiscounted credit purchases. Terms | the first 180 days, 60-day cookie | https://api.novita.ai/openai/v1/models |
| OpenRouter | Nothing | — | https://openrouter.ai/api/v1/models |
Considered and left out, each for a reason that is checked again by 2026-12-22:
The programme terms were read from each provider on 2026-09-22. Who runs this site.