Rig reports. Nothing estimated.
If you want to buy hardware with your eyes open, the thing worth knowing is what it can actually do. Each report is a machine we rented, with every model loaded and timed on it.
- FREE →NVIDIA GeForce RTX 4090289 tok/s on gpt-oss 20B · measured 2026-08-01
- FREE →NVIDIA GeForce RTX 3090227 tok/s on gpt-oss 20B · measured 2026-07-31
- PRO · $25 →1x L40S 48GB231 tok/s on gpt-oss 20B · measured 2026-08-30
- PRO · $25 →1x RTX A6000 48GB206 tok/s on gpt-oss 20B · measured 2026-08-30
- PRO · $25 →2x A40 48GB180 tok/s on gpt-oss 20B · measured 2026-08-30
- LIGHT · $9 →NVIDIA GeForce RTX 3060 Ti123 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
- LIGHT · $9 →NVIDIA GeForce RTX 3070126 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
- PRO · $25 →NVIDIA GeForce RTX 5060 Ti 16 GB150 tok/s on gpt-oss 20B · measured 2026-08-29
- LIGHT · $9 →NVIDIA GeForce RTX 4070151 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
- LIGHT · $9 →NVIDIA GeForce RTX 5070186 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
- LIGHT · $9 →NVIDIA GeForce RTX 3060 12 GB98 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
- PRO · $25 →1x H100 SXM 80GB278 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
- PRO · $25 →2x L40S144 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
- PRO · $25 →1x H200 141GB280 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
- PRO · $25 →NVIDIA RTX 4000 Ada 20GB127 tok/s on Qwen3 30B A3B Instruct 2507 · measured 2026-08-25
- PRO · $25 →1x NVIDIA L492 tok/s on gpt-oss 20B · measured 2026-08-25
- PRO · $25 →1x NVIDIA A40182 tok/s on gpt-oss 20B · measured 2026-08-25
- PRO · $25 →NVIDIA GeForce RTX 5090427 tok/s on gpt-oss 20B · measured 2026-08-21
- COMMISSION IT →Apple M1 (8GB)
- COMMISSION IT →Apple M4 (16GB)
- COMMISSION IT →AMD Radeon RX 9070 XT
- COMMISSION IT →Apple M1 Pro (16GB)
- COMMISSION IT →Apple M4 Pro (24GB)
- COMMISSION IT →NVIDIA Quadro P3200 Mobile
- COMMISSION IT →NVIDIA GeForce RTX 5070 Ti
- COMMISSION IT →Apple M3 Pro (18GB)
- COMMISSION IT →NVIDIA GeForce RTX 5080
- COMMISSION IT →NVIDIA GeForce RTX 3080 Ti
- COMMISSION IT →Apple M3 Max (36GB)
- COMMISSION IT →NVIDIA GeForce RTX 3050 Mobile
These machines were rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.
You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.
ASK ABOUT A REPORT →WHAT WE MEASURE, ON EVERY RIG7,349 FIGURES ACROSS 18 MACHINES
Tokens per second generating with the model fully resident on the GPU. Median of several independent runs, each from a cold load.
Tokens per second processing the prompt — the half of latency that a decode figure does not describe.
Time to first token at the median and the 95th percentile. p95 is what an interactive client actually experiences.
The maximum the driver reported during the run, not the weight size. The gap between them is the runtime’s own allocation, and it is why a model fits on paper and not on the card.
The same weights at several rungs on the same card: what each extra bit per weight costs in tokens per second and in VRAM.
Speed with the KV cache already filled to 4K, 8K, 16K, 32K and beyond — not the empty-cache figure, and with the depth at which the model stops loading.
Tokens per second at each fraction of the model kept resident, for the case where it does not fit and layers cross PCIe every token.
Aggregate throughput, per-stream throughput and TTFT p95 at 1, 2, 4, 8 and 16 streams, measured over the seconds the server was actually saturated.
Deterministic probes at temperature 0 with a programmatic check, plus a separate looping test — because a damaged quantisation posts a higher token rate, not a lower one.
Watts sampled at the card during generation, and cost per million tokens at the rate the machine was billed, single stream and batched.
On and off, at depth. Both LM Studio and Ollama enable it by default; whether it helps depends on the model and the cache size.
The card’s decode constants and implied bandwidth, the fixed VRAM overhead before any weights, and how TTFT scales with concurrency — each with its n and its correlation.
Every speed is the median of several independent processes, each of which reloaded the model from cold; the runs and their spread are printed. Models that failed to load are listed with the stage they failed at, and each report carries the raw JSON it was rendered from.
September 2026 — what rent did.
Computed, not written. Every figure comes out of the rates we record from RunPod, Vast.ai and DataCrunch each morning — which means it can be regenerated for any past month and come out the same, and that it cannot exist for a month we did not watch. Free, like the rest of this page except the measurements above.
September 2026 has 7 days recorded, across 103 cards. A report needs 14.
Below that, one provider outage or a single odd weekend moves every figure in the document and nothing on the page would show it. The record grows by a day each morning, and this becomes a report on its own.
Every VRAM calculator on the internet runs the same roofline we do and tells you to go validate it yourself. We went and validated it. The measurements here are also what corrects the estimator: the context curve, the offload curve and the KV-cache formula on this site were all rewritten from these files after they disagreed with the machine. How the estimates work → What a commissioned report contains →