FitMyLLM
▸ MEASURED RIG REPORT$9LIGHT · 7 MODELS

1x RTX 3070

7 models measured8 GB VRAMllama.cpp 91f8c9c2026-08-290.6 h on the machine

Before you spend $325 on a second-hand one, find out if it does what you need.

7 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ WHAT THIS REPORT SETTLES
Does Qwen3 8B run on it?
It did, at Q6_K — 6.6 GB of 8 GB at peak
And how fast, model by model?
Decode, prefill and TTFT p50/p95 on 7 models
Which quantisation is worth running?
1 model run across several rungs; Q4_K_M, Q3_K_M, Q5_K_M, Q6_K measured on this card
How long a prompt before it slows down, or stops fitting?
Decode and peak VRAM at 4K → 32K, on 4 models
And when a model is bigger than the card?
tok/s at each resident fraction, 4 models spilled into system RAM
How many people can use it at once?
1 → 8 streams: aggregate tok/s, per-stream tok/s and TTFT p95
Does the quantisation cost answers?
Probed at temperature 0 on 7 models, scored per category
What does it cost to keep running?
Watts at the card, and dollars per million tokens, single stream and batched
§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
Qwen3 4B Instruct 2507 · Q4_K_M126
Llama 3.1 8B Instruct · Q4_K_M80.9
Qwen3 8B · Q4_K_M78.4
Qwen3 8B (Q3_K_M)73.3
Qwen3.5 9B · Q4_K_M69.8
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 2 models, and every other measurement, are listed below.

What is in this report

MODELS RUN ON THIS MACHINE7 MEASURED · 2 FAILED
  • · Qwen3 4B Instruct 2507 · Q4_K_M
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Qwen3 8B · Q4_K_M
  • · Qwen3.5 9B · Q4_K_M
  • · Qwen3 8B (Q3_K_M)
  • · Qwen3 8B (Q5_K_M)
  • · Qwen3 8B (Q6_K)
  • × Gemma 3 12B Instruct · Q4_K_M
  • × Qwen3 14B · Q4_K_M

Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.

EVERY FIELD, MEASUREMENT BY MEASUREMENT
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K
  • · 1 model run across several rungs on this card
CONTEXT DEPTH
  • · Decode measured with the cache filled to 4K, 8K, 16K, 32K
  • · 4 models swept, with peak VRAM at each depth
OFFLOAD TO SYSTEM RAM
  • · 4 models run at several resident fractions
  • · tok/s at each fraction, against the fully resident rate
CONCURRENCY
  • · Streams: 1, 2, 4, 8
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 7 models load-tested
ACCURACY AFTER QUANTISATION
  • · 7 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 7 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA GeForce RTX 3070
  • · llama.cpp build 91f8c9c5fb038c086e13e9cd823c29b33b07ba54
  • · Driver 595.71.05
  • · Run finished 2026-08-29
  • · 0.6 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.