▸ CLOUD GPU PRICESLIVEprices read seconds ago
What a million tokens actually costs.
Every provider publishes an hourly rate. None of them publish what it buys — and an hour of a dearer, faster card is often cheaper than an hour of a slow one. Converting one into the other takes a throughput figure, so that is the column we add.
RunPod36 ratesVast.ai50 ratesDataCrunch32 rates
MODEL
§ 01 · EVERY QUANTISATION, PRICEDCLICK A ROW TO COMPARE IT
Each bar is the cheapest cost per million tokens at that size, on the cheapest card that holds it. A heavier quantisation costs more per token twice over — it is slower, and it needs a bigger card.
CHEAPEST
$0.160
per 1M · Vast.ai · RTX 3090
MEASURED
15
of 63 rows · best measured $0.160/1M
FASTEST
427 tok/s
RTX 5090 · $0.263/1M
OPTIONS
30
cards · 63 quotes · 15 measured
§ 02 · WHERE TO RUN IT AT Q4_K_M63 OPTIONS
| GPU | VRAM | PROVIDER | PER HOUR | TOK/S | PER 1M TOKENS ▾ | 30-DAY |
|---|---|---|---|---|---|---|
| RTX 3090BEST | 24 GB | Vast.aicheapest listing of 2 | $0.131 | 227MEAS | $0.160 | 4D |
| RTX 4070 Ti SUPER | 16 GB | Vast.aicheapest listing of 1 | $0.118 | 141EST | $0.232 | 4D |
| RTX 3090 Ti | 24 GB | Vast.aicheapest listing of 1 | $0.169 | 197EST | $0.238 | 3D |
| RTX 5060 Ti 16GB | 16 GB | Vast.aicheapest listing of 2 | $0.089 | 102EST | $0.241 | 4D |
| RTX 4080 SUPER | 16 GB | Vast.aicheapest listing of 1 | $0.138 | 151EST | $0.253 | 4D |
| RTX 5080 | 16 GB | Vast.aicheapest listing of 1 | $0.178 | 189EST | $0.261 | 4D |
| RTX 5090 | 32 GB | Vast.aicheapest listing of 2 | $0.404 | 427MEAS | $0.263 | 4D |
| RTX 3090 | 24 GB | RunPodpublished rate | $0.220 | 227MEAS | $0.269 | 4D |
| TITAN RTX | 24 GB | Vast.aicheapest listing of 1 | $0.149 | 141EST | $0.293 | 2D |
| Tesla V100 SXM2 16 GB | 16 GB | RunPodpublished rate | $0.230 | 216EST | $0.296 | 4D |
| RTX 4090 | 24 GB | Vast.aicheapest listing of 2 | $0.339 | 289MEAS | $0.326 | 3D |
| RTX 4090 | 24 GB | RunPodpublished rate | $0.340 | 289MEAS | $0.327 | 4D |
| RTX A4500 | 20 GB | RunPodpublished rate | $0.190 | 135EST | $0.391 | 4D |
| RTX A6000 48GB | 48 GB | DataCrunchinterruptible | $0.305 | 195MEAS | $0.434 | 4D |
| RTX 5090 | 32 GB | RunPodpublished rate | $0.690 | 427MEAS | $0.449 | 4D |
| RTX A6000 48GB | 48 GB | RunPodpublished rate | $0.330 | 195MEAS | $0.470 | 4D |
| RTX 4080 | 16 GB | Vast.aicheapest listing of 1 | $0.269 | 148EST | $0.505 | 4D |
| RTX 4080 SUPER | 16 GB | RunPodpublished rate | $0.280 | 151EST | $0.515 | 4D |
| A40 48GB | 48 GB | RunPodpublished rate | $0.350 | 182MEAS | $0.534 | 4D |
| RTX 5080 | 16 GB | RunPodpublished rate | $0.390 | 189EST | $0.573 | 4D |
▸ WHAT THESE NUMBERS ARE
- Rates are read live from each provider’s own API — refreshed on this page every three minutes, cached ten minutes at the edge. Last read seconds ago.
- MEAS is a speed we measured on that card with that model, from a rig report. EST is our engine’s prediction — a cost built on it inherits the engine’s error.
- Cost per million tokens is the hourly rate divided by output tokens at one stream. Batched serving is cheaper per token; an idle card still bills.
- Datacenter cards rank badly here, and that is the finding, not a fault. An H100 or H200 earns its price by serving many streams at once; at one stream its enormous bandwidth sits mostly idle, and our engine models that as a much lower efficiency than a consumer card gets. If you are serving a crowd rather than yourself, this ranking inverts — that case is the enterprise planner, not this page. Those rows are also the least certain on the page: they are estimates, and the constant behind them is the one we most want to replace with a measurement.
- A marketplace price is one host’s listing, not a published rate: it can vanish, and hosts may run a card below its stock power limit, which costs speed the price does not show.
- Only cards holding the model at the chosen quantisation, with 10% left for the runtime, are listed. Nothing here is priced through CPU offload.
Provider links may carry a referral code. It costs you nothing, it does not change what is listed, and it never changes the order — the table sorts by the column you choose and by nothing else. See methodology.