FitMyLLM
▸ MEASURED RIG REPORT$25PRO · 14 MODELS

1x RTX 5060 Ti 16GB

14 models measured16 GB VRAMllama.cpp 91f8c9c2026-08-292.4 h on the machine

Before you spend $429 on this card, find out if it does what you need.

14 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ WHAT THIS REPORT SETTLES
Does Mistral Small 3.2 24B Instruct run on it?
It did, at Q4_K_M — 14.0 GB of 16 GB at peak
And how fast, model by model?
Decode, prefill and TTFT p50/p95 on 14 models
Which quantisation is worth running?
2 models run across several rungs; Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0 measured on this card
How long a prompt before it slows down, or stops fitting?
Decode and peak VRAM at 4K → 32K, on 8 models
And when a model is bigger than the card?
tok/s at each resident fraction, 8 models spilled into system RAM
How many people can use it at once?
1 → 16 streams: aggregate tok/s, per-stream tok/s and TTFT p95
Does the quantisation cost answers?
Probed at temperature 0 on 14 models, scored per category
What does it cost to keep running?
Watts at the card, and dollars per million tokens, single stream and batched
§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
gpt-oss 20B · Q4_K_M150
Qwen3 4B Instruct 2507 · Q4_K_M136
Llama 3.1 8B Instruct · Q4_K_M84.5
Qwen3 8B (Q3_K_M)83.1
Qwen3 8B · Q4_K_M82.5
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 9 models, and every other measurement, are listed below.

What is in this report

MODELS RUN ON THIS MACHINE14 MEASURED · 3 FAILED
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Qwen3 4B Instruct 2507 · Q4_K_M
  • · Qwen3 8B · Q4_K_M
  • · Qwen3.5 9B · Q4_K_M
  • · Gemma 3 12B Instruct · Q4_K_M
  • · Qwen3 14B · Q4_K_M
  • · gpt-oss 20B · Q4_K_M
  • · Mistral Small 3.2 24B Instruct · Q4_K_M
  • × Gemma 3 27B Instruct · Q4_K_M
  • × Qwen3 30B A3B Instruct 2507 · Q4_K_M
  • × GLM 4.7 Flash · Q4_K_M
  • · Qwen3 8B (Q3_K_M)
  • · Qwen3 8B (Q5_K_M)
  • · Qwen3 8B (Q6_K)
  • · Qwen3 8B (Q8_0)
  • · Gemma 3 12B Instruct (Q5_K_M)
  • · Gemma 3 12B Instruct (Q8_0)

Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.

EVERY FIELD, MEASUREMENT BY MEASUREMENT
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K, Q8_0
  • · 2 models run across several rungs on this card
CONTEXT DEPTH
  • · Decode measured with the cache filled to 4K, 8K, 16K, 32K
  • · 8 models swept, with peak VRAM at each depth
OFFLOAD TO SYSTEM RAM
  • · 8 models run at several resident fractions
  • · tok/s at each fraction, against the fully resident rate
CONCURRENCY
  • · Streams: 1, 2, 4, 8, 16
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 14 models load-tested
ACCURACY AFTER QUANTISATION
  • · 14 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 14 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA GeForce RTX 5060 Ti
  • · llama.cpp build 91f8c9c5fb038c086e13e9cd823c29b33b07ba54
  • · Driver 610.43.02
  • · Run finished 2026-08-29
  • · 2.4 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.