1x RTX 5090
17 models were loaded on this machine and timed: decode, prefill, TTFT p50 and p95, peak VRAM, load time, and where the run covered them, quantisation ladders, context depth, RAM offload, concurrency, accuracy probes and power. The llama.cpp build and the driver version are recorded with every figure. Full contents below.
Two reports of the same shape are open in full, if you want to read one before paying: the RTX 4090 and the RTX 3090.
One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.
Decode only, single stream, for the fastest 5. The remaining 12 models, and every other measurement, are listed below.
What is in this report
One run on one machine. Everything below was measured on the hardware named above, in a single session, on one llama.cpp build.
- · Llama 3.1 8B Instruct · Q4_K_M
- · Qwen3 4B Instruct 2507 · Q4_K_M
- · Qwen3 8B · Q4_K_M
- · Qwen3.5 9B · Q4_K_M
- · Gemma 3 12B Instruct · Q4_K_M
- · Qwen3 14B · Q4_K_M
- · gpt-oss 20B · Q4_K_M
- · Mistral Small 3.2 24B Instruct · Q4_K_M
- · Gemma 3 27B Instruct · Q4_K_M
- · Qwen3 30B A3B Instruct 2507 · Q4_K_M
- · GLM 4.7 Flash · Q4_K_M
- · Qwen3 8B (Q3_K_M)
- · Qwen3 8B (Q5_K_M)
- · Qwen3 8B (Q6_K)
- · Qwen3 8B (Q8_0)
- · Gemma 3 12B Instruct (Q5_K_M)
- · Gemma 3 12B Instruct (Q8_0)
- · Decode, single stream, tok/s
- · Prefill, tok/s
- · TTFT p50 and p95, ms
- · Peak VRAM reported by the driver
- · Weight size on disk, and load time
- · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0
- · 2 models run across several rungs on this card
- · Decode measured with the cache filled to 4K, 16K
- · 11 models swept, with peak VRAM at each depth
- · 11 models run at several resident fractions
- · tok/s at each fraction, against the fully resident rate
- · Streams: 1, 2, 4, 8, 16
- · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
- · 17 models load-tested
- · 17 models probed at temperature 0
- · Score per category, and a separate looping check
- · Watts sampled at the card for 17 models
- · Cost per million tokens, single stream and batched
- · The card's decode constants: fixed ms per token, and implied bandwidth
- · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
- · TTFT p95 against concurrent streams
- · Each with its n and its correlation; a weak fit is quoted without a line
- · GPU: NVIDIA GeForce RTX 5090
- · llama.cpp build b10156
- · Driver 580.126.09
- · Run finished 2026-08-21
- · 4.0 hours on the machine
- · This page, in full, at this URL
- · A PDF with the same tables and charts
- · The underlying JSON, through the API, once the report is yours
- · No expiry, and no subscription to keep reading it
Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.