FitMyLLM
▸ MEASURED RIG REPORT$25PRO · 11 MODELS

1x NVIDIA L4

11 models measured22 GB VRAMllama.cpp b101562026-08-252.8 h on the machine

Before you spend $2,500 on this card, find out if it does what you need.

11 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ WHAT THIS REPORT SETTLES
Does Mistral Small 3.2 24B Instruct run on it?
It did, at Q4_K_M — 14.1 GB of 22 GB at peak
And how fast, model by model?
Decode, prefill and TTFT p50/p95 on 11 models
Which quantisation is worth running?
2 models run across several rungs; Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0 measured on this card
And when a model is bigger than the card?
tok/s at each resident fraction, 8 models spilled into system RAM
How many people can use it at once?
1 → 8 streams: aggregate tok/s, per-stream tok/s and TTFT p95
Does the quantisation cost answers?
Probed at temperature 0 on 11 models, scored per category
What does it cost to keep running?
Watts at the card, and dollars per million tokens, single stream and batched
§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
gpt-oss 20B · Q4_K_M92.0
Qwen3 4B Instruct 2507 · Q4_K_M83.3
Llama 3.1 8B Instruct · Q4_K_M50.3
Qwen3 8B · Q4_K_M48.9
Qwen3.5 9B · Q4_K_M43.5
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 6 models, and every other measurement, are listed below.

What is in this report

MODELS RUN ON THIS MACHINE11 MEASURED · 6 FAILED
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Qwen3 4B Instruct 2507 · Q4_K_M
  • · Qwen3 8B · Q4_K_M
  • · Qwen3.5 9B · Q4_K_M
  • · Gemma 3 12B Instruct · Q4_K_M
  • · Qwen3 14B · Q4_K_M
  • · gpt-oss 20B · Q4_K_M
  • · Mistral Small 3.2 24B Instruct · Q4_K_M
  • × Gemma 3 27B Instruct · Q4_K_M
  • × Qwen3 30B A3B Instruct 2507 · Q4_K_M
  • × GLM 4.7 Flash · Q4_K_M
  • × Qwen3 8B (Q3_K_M)
  • × Qwen3 8B (Q5_K_M)
  • × Qwen3 8B (Q6_K)
  • · Qwen3 8B (Q8_0)
  • · Gemma 3 12B Instruct (Q5_K_M)
  • · Gemma 3 12B Instruct (Q8_0)

Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.

EVERY FIELD, MEASUREMENT BY MEASUREMENT
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0
  • · 2 models run across several rungs on this card
OFFLOAD TO SYSTEM RAM
  • · 8 models run at several resident fractions
  • · tok/s at each fraction, against the fully resident rate
CONCURRENCY
  • · Streams: 1, 2, 4, 8
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 11 models load-tested
ACCURACY AFTER QUANTISATION
  • · 11 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 11 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA L4
  • · llama.cpp build b10156
  • · Driver 570.195.03
  • · Run finished 2026-08-25
  • · 2.8 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.