Buy a GPU, rent one, or pay by the token?
Pick a model and put in how you would use it. The cards that hold it and their speed come from the engine, the rent and per-token prices are read live from the providers, and the prices you know better — the card you would buy, your electricity — are yours to enter.
MiMo-V2-Flash at Q4_K_M: a year of it, three ways
Prices · 18:30 UTC, 22 SeptWatts start at the Apple M5 Ultra’s published board power, an upper bound. No card price is filled in: street prices moved too fast in 2026 for a default to be fair. How it runs on this card →
Cheapest at this usage: rent 8× L40S, about $5,452.81 a year over three years.
- Per-token is cheapest for light, bursty use; a machine wins when it stays busy, runs a model no provider serves, or data must stay on it.
Referral links Vast.ai, RunPod and Novita pay us a share of what you spend if you sign up through these buttons. It costs you nothing, and it never decides an order or a recommendation: both are computed from the live price and the speed, and options that pay us nothing are listed and recommended on the same terms. How we rank
Find the Apple M5 Ultra: Amazon ↗ · eBay (new and used) ↗Store links may pay us a commission. They never decide which card is suggested — the memory arithmetic does.
How to read it
- Per-token wins for light use. A provider batches your requests with everyone else’s, so one person chatting pays a fraction of an hour of a whole GPU.
- A machine wins when it stays busy — long agent runs, batch jobs, many users — or when the model is a fine-tune nobody serves, or the data must not leave a machine you control.
- Owning wins with hours. The break-even line says how many years at your hours; it moves a lot with the electricity price and with what the card resells for.
- Speeds are single-stream estimates from the engine, and the machines are sized at Q4_K_M with 8,192 tokens of context. For every format and every card, see the MiMo-V2-Flash page.