FitMyLLM
▸ MEASURED RIG REPORT€24

1x RTX A6000 (MoE ladder)

12 models measured48 GB VRAMllama.cpp b101562026-08-011.9 h on the machine

12 models were loaded on this machine and timed: decode, prefill, TTFT p50 and p95, peak VRAM, load time, and where the run covered them, quantisation ladders, context depth, RAM offload, concurrency, accuracy probes and power. The llama.cpp build and the driver version are recorded with every figure. Full contents below.

Two reports of the same shape are open in full, if you want to read one before paying: the RTX 4090 and the RTX 3090.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
LFM2 8B A1B · Q4_K_M391
Ling lite 16.8B · Q4_K_M238
Granite 4.0 H Tiny · Q4_K_M222
DeepSeek Coder V2 Lite 16B · Q4_K_M210
gpt-oss 20B · Q4_K_M195
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 7 models, and every other measurement, are listed below.

What is in this report

One run on one machine. Everything below was measured on the hardware named above, in a single session, on one llama.cpp build.

MODELS RUN ON THIS MACHINE12 MEASURED
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Granite 4.0 H Tiny · Q4_K_M
  • · LFM2 8B A1B · Q4_K_M
  • · gpt-oss 20B · Q4_K_M
  • · Phi-3.5 MoE 42B · Q4_K_M
  • · DeepSeek Coder V2 Lite 16B · Q4_K_M
  • · Ling lite 16.8B · Q4_K_M
  • · ERNIE 4.5 21B A3B · Q4_K_M
  • · Qwen3 30B A3B Instruct 2507 · Q4_K_M
  • · Qwen3 30B A3B Instruct 2507 (Q3_K_M)
  • · Qwen3 30B A3B Instruct 2507 (Q8_0)
  • · GLM 4.7 Flash · Q4_K_M
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M, Q3_K_M, Q8_0
  • · 1 model run across several rungs on this card
CONTEXT DEPTH
  • · Decode measured with the cache filled to 4K, 16K
  • · 10 models swept, with peak VRAM at each depth
CONCURRENCY
  • · Streams: 1, 4, 16
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 12 models load-tested
ACCURACY AFTER QUANTISATION
  • · 12 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 12 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA RTX A6000
  • · llama.cpp build b10156
  • · Driver 570.195.03
  • · Run finished 2026-08-01
  • · 1.9 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.