1x RTX A6000 (MoE ladder)
12 models were loaded on this machine and timed: decode, prefill, TTFT p50 and p95, peak VRAM, load time, and where the run covered them, quantisation ladders, context depth, RAM offload, concurrency, accuracy probes and power. The llama.cpp build and the driver version are recorded with every figure. Full contents below.
Two reports of the same shape are open in full, if you want to read one before paying: the RTX 4090 and the RTX 3090.
One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.
Decode only, single stream, for the fastest 5. The remaining 7 models, and every other measurement, are listed below.
What is in this report
One run on one machine. Everything below was measured on the hardware named above, in a single session, on one llama.cpp build.
- · Llama 3.1 8B Instruct · Q4_K_M
- · Granite 4.0 H Tiny · Q4_K_M
- · LFM2 8B A1B · Q4_K_M
- · gpt-oss 20B · Q4_K_M
- · Phi-3.5 MoE 42B · Q4_K_M
- · DeepSeek Coder V2 Lite 16B · Q4_K_M
- · Ling lite 16.8B · Q4_K_M
- · ERNIE 4.5 21B A3B · Q4_K_M
- · Qwen3 30B A3B Instruct 2507 · Q4_K_M
- · Qwen3 30B A3B Instruct 2507 (Q3_K_M)
- · Qwen3 30B A3B Instruct 2507 (Q8_0)
- · GLM 4.7 Flash · Q4_K_M
- · Decode, single stream, tok/s
- · Prefill, tok/s
- · TTFT p50 and p95, ms
- · Peak VRAM reported by the driver
- · Weight size on disk, and load time
- · Rungs measured: Q4_K_M, Q3_K_M, Q8_0
- · 1 model run across several rungs on this card
- · Decode measured with the cache filled to 4K, 16K
- · 10 models swept, with peak VRAM at each depth
- · Streams: 1, 4, 16
- · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
- · 12 models load-tested
- · 12 models probed at temperature 0
- · Score per category, and a separate looping check
- · Watts sampled at the card for 12 models
- · Cost per million tokens, single stream and batched
- · The card's decode constants: fixed ms per token, and implied bandwidth
- · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
- · TTFT p95 against concurrent streams
- · Each with its n and its correlation; a weak fit is quoted without a line
- · GPU: NVIDIA RTX A6000
- · llama.cpp build b10156
- · Driver 570.195.03
- · Run finished 2026-08-01
- · 1.9 hours on the machine
- · This page, in full, at this URL
- · A PDF with the same tables and charts
- · The underlying JSON, through the API, once the report is yours
- · No expiry, and no subscription to keep reading it
Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.