FitMyLLM
▸ MEASURED RIG REPORT€24

1x NVIDIA L40S

17 models measured45 GB VRAMllama.cpp b101562026-08-250.9 h on the machine

17 models were loaded on this machine and timed: decode, prefill, TTFT p50 and p95, peak VRAM, load time, and where the run covered them, quantisation ladders, context depth, RAM offload, concurrency, accuracy probes and power. The llama.cpp build and the driver version are recorded with every figure. Full contents below.

Two reports of the same shape are open in full, if you want to read one before paying: the RTX 4090 and the RTX 3090.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
gpt-oss 20B · Q4_K_M234
Qwen3 30B A3B Instruct 2507 · Q4_K_M220
Qwen3 4B Instruct 2507 · Q4_K_M212
GLM 4.7 Flash · Q4_K_M158
Qwen3 8B (Q3_K_M)150
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 12 models, and every other measurement, are listed below.

What is in this report

One run on one machine. Everything below was measured on the hardware named above, in a single session, on one llama.cpp build.

MODELS RUN ON THIS MACHINE17 MEASURED
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Qwen3 4B Instruct 2507 · Q4_K_M
  • · Qwen3 8B · Q4_K_M
  • · Qwen3.5 9B · Q4_K_M
  • · Gemma 3 12B Instruct · Q4_K_M
  • · Qwen3 14B · Q4_K_M
  • · gpt-oss 20B · Q4_K_M
  • · Mistral Small 3.2 24B Instruct · Q4_K_M
  • · Gemma 3 27B Instruct · Q4_K_M
  • · Qwen3 30B A3B Instruct 2507 · Q4_K_M
  • · GLM 4.7 Flash · Q4_K_M
  • · Qwen3 8B (Q3_K_M)
  • · Qwen3 8B (Q5_K_M)
  • · Qwen3 8B (Q6_K)
  • · Qwen3 8B (Q8_0)
  • · Gemma 3 12B Instruct (Q5_K_M)
  • · Gemma 3 12B Instruct (Q8_0)
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0
  • · 2 models run across several rungs on this card
OFFLOAD TO SYSTEM RAM
  • · 11 models run at several resident fractions
  • · tok/s at each fraction, against the fully resident rate
CONCURRENCY
  • · Streams: 1, 2, 4, 8
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 17 models load-tested
ACCURACY AFTER QUANTISATION
  • · 17 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 17 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA L40S
  • · llama.cpp build b10156
  • · Driver 580.126.09
  • · Run finished 2026-08-25
  • · 0.9 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.