▸ MEASURED RIG REPORT$25PRO · 17 MODELS
1x RTX 5090
17 models measured32 GB VRAMllama.cpp b101562026-08-214.0 h on the machine
Before you spend $1,999 on this card, find out if it does what you need.
17 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.
One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.
§ WHAT THIS REPORT SETTLES
- Does Qwen3 30B A3B Instruct 2507 run on it?
- It did, at Q4_K_M — 18.1 GB of 32 GB at peak
- And how fast, model by model?
- Decode, prefill and TTFT p50/p95 on 17 models
- Which quantisation is worth running?
- 2 models run across several rungs; Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0 measured on this card
- How long a prompt before it slows down, or stops fitting?
- Decode and peak VRAM at 4K → 16K, on 11 models
- And when a model is bigger than the card?
- tok/s at each resident fraction, 11 models spilled into system RAM
- How many people can use it at once?
- 1 → 16 streams: aggregate tok/s, per-stream tok/s and TTFT p95
- Does the quantisation cost answers?
- Probed at temperature 0 on 17 models, scored per category
- What does it cost to keep running?
- Watts at the card, and dollars per million tokens, single stream and batched
§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
427
373
353
267
260
DECODE TOK/S, SINGLE STREAM
Decode only, single stream, for the fastest 5. The remaining 12 models, and every other measurement, are listed below.
What is in this report
MODELS RUN ON THIS MACHINE17 MEASURED
- · Llama 3.1 8B Instruct · Q4_K_M
- · Qwen3 4B Instruct 2507 · Q4_K_M
- · Qwen3 8B · Q4_K_M
- · Qwen3.5 9B · Q4_K_M
- · Gemma 3 12B Instruct · Q4_K_M
- · Qwen3 14B · Q4_K_M
- · gpt-oss 20B · Q4_K_M
- · Mistral Small 3.2 24B Instruct · Q4_K_M
- · Gemma 3 27B Instruct · Q4_K_M
- · Qwen3 30B A3B Instruct 2507 · Q4_K_M
- · GLM 4.7 Flash · Q4_K_M
- · Qwen3 8B (Q3_K_M)
- · Qwen3 8B (Q5_K_M)
- · Qwen3 8B (Q6_K)
- · Qwen3 8B (Q8_0)
- · Gemma 3 12B Instruct (Q5_K_M)
- · Gemma 3 12B Instruct (Q8_0)
EVERY FIELD, MEASUREMENT BY MEASUREMENT
FOR EVERY MODEL
- · Decode, single stream, tok/s
- · Prefill, tok/s
- · TTFT p50 and p95, ms
- · Peak VRAM reported by the driver
- · Weight size on disk, and load time
QUANTISATION
- · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0
- · 2 models run across several rungs on this card
CONTEXT DEPTH
- · Decode measured with the cache filled to 4K, 16K
- · 11 models swept, with peak VRAM at each depth
OFFLOAD TO SYSTEM RAM
- · 11 models run at several resident fractions
- · tok/s at each fraction, against the fully resident rate
CONCURRENCY
- · Streams: 1, 2, 4, 8, 16
- · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
- · 17 models load-tested
ACCURACY AFTER QUANTISATION
- · 17 models probed at temperature 0
- · Score per category, and a separate looping check
POWER AND COST
- · Watts sampled at the card for 17 models
- · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
- · The card's decode constants: fixed ms per token, and implied bandwidth
- · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
- · TTFT p95 against concurrent streams
- · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
- · GPU: NVIDIA GeForce RTX 5090
- · llama.cpp build b10156
- · Driver 580.126.09
- · Run finished 2026-08-21
- · 4.0 hours on the machine
WHAT YOU GET
- · This page, in full, at this URL
- · A PDF with the same tables and charts
- · The underlying JSON, through the API, once the report is yours
- · No expiry, and no subscription to keep reading it
Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.