Buy a GPU, rent one, or pay by the token?
Pick a model and put in how you would use it. The cards that hold it and their speed come from the engine, the rent and per-token prices are read live from the providers, and the prices you know better — the card you would buy, your electricity — are yours to enter.
gpt-neox-20b at Q4_K_M: a year of it, three ways
Prices · 18:30 UTC, 22 SeptNo single card a person can buy holds gpt-neox-20b at Q4_K_M; owning means a multi-GPU machine.
Cheapest at this usage: rent RTX 3090, about $178.41 a year over three years.
- Per-token is cheapest for light, bursty use; a machine wins when it stays busy, runs a model no provider serves, or data must stay on it.
Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
How to read it
- Per-token wins for light use. A provider batches your requests with everyone else’s, so one person chatting pays a fraction of an hour of a whole GPU.
- A machine wins when it stays busy — long agent runs, batch jobs, many users — or when the model is a fine-tune nobody serves, or the data must not leave a machine you control.
- Owning wins with hours. The break-even line says how many years at your hours; it moves a lot with the electricity price and with what the card resells for.
- Speeds are single-stream estimates from the engine, and the machines are sized at Q4_K_M with 8,192 tokens of context. For every format and every card, see the gpt-neox-20b page.