1x RTX 3070
Before you spend $325 on a second-hand one, find out if it does what you need.
7 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.
One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.
- Does Qwen3 8B run on it?
- It did, at Q6_K — 6.6 GB of 8 GB at peak
- And how fast, model by model?
- Decode, prefill and TTFT p50/p95 on 7 models
- Which quantisation is worth running?
- 1 model run across several rungs; Q4_K_M, Q3_K_M, Q5_K_M, Q6_K measured on this card
- How long a prompt before it slows down, or stops fitting?
- Decode and peak VRAM at 4K → 32K, on 4 models
- And when a model is bigger than the card?
- tok/s at each resident fraction, 4 models spilled into system RAM
- How many people can use it at once?
- 1 → 8 streams: aggregate tok/s, per-stream tok/s and TTFT p95
- Does the quantisation cost answers?
- Probed at temperature 0 on 7 models, scored per category
- What does it cost to keep running?
- Watts at the card, and dollars per million tokens, single stream and batched
Decode only, single stream, for the fastest 5. The remaining 2 models, and every other measurement, are listed below.
What is in this report
- · Qwen3 4B Instruct 2507 · Q4_K_M
- · Llama 3.1 8B Instruct · Q4_K_M
- · Qwen3 8B · Q4_K_M
- · Qwen3.5 9B · Q4_K_M
- · Qwen3 8B (Q3_K_M)
- · Qwen3 8B (Q5_K_M)
- · Qwen3 8B (Q6_K)
- × Gemma 3 12B Instruct · Q4_K_M
- × Qwen3 14B · Q4_K_M
Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.
EVERY FIELD, MEASUREMENT BY MEASUREMENT
- · Decode, single stream, tok/s
- · Prefill, tok/s
- · TTFT p50 and p95, ms
- · Peak VRAM reported by the driver
- · Weight size on disk, and load time
- · Rungs measured: Q4_K_M, Q3_K_M, Q5_K_M, Q6_K
- · 1 model run across several rungs on this card
- · Decode measured with the cache filled to 4K, 8K, 16K, 32K
- · 4 models swept, with peak VRAM at each depth
- · 4 models run at several resident fractions
- · tok/s at each fraction, against the fully resident rate
- · Streams: 1, 2, 4, 8
- · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
- · 7 models load-tested
- · 7 models probed at temperature 0
- · Score per category, and a separate looping check
- · Watts sampled at the card for 7 models
- · Cost per million tokens, single stream and batched
- · The card's decode constants: fixed ms per token, and implied bandwidth
- · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
- · TTFT p95 against concurrent streams
- · Each with its n and its correlation; a weak fit is quoted without a line
- · GPU: NVIDIA GeForce RTX 3070
- · llama.cpp build 91f8c9c5fb038c086e13e9cd823c29b33b07ba54
- · Driver 595.71.05
- · Run finished 2026-08-29
- · 0.6 hours on the machine
- · This page, in full, at this URL
- · A PDF with the same tables and charts
- · The underlying JSON, through the API, once the report is yours
- · No expiry, and no subscription to keep reading it
Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.