Everything below is also a PDF — every table, every chart, the machine and the method, in one document you can keep or forward. The raw JSON is the same measurements as data, for anyone who would rather check them than read them.
Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.
Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIX
What runs, and how fast
Decode: single-stream generation, model fully resident on the GPU. Peak VRAM: the maximum the driver reported during the run, not the weight size — the difference is the runtime's own allocation. TTFT p95: time to first token at the 95th percentile, measured under the load stated in the concurrency table.
gpt-oss 20B · Q4_K_M289
Qwen3 30B A3B Instruct 2507 · Q4_K_M259
Qwen3 4B Instruct 2507 · Q4_K_M258
Qwen3 8B (Q3_K_M)187
GLM 4.7 Flash · Q4_K_M186
Llama 3.1 8B Instruct · Q4_K_M171
Qwen3 8B · Q4_K_M163
Qwen3 8B (Q5_K_M)145
Qwen3.5 9B · Q4_K_M144
Qwen3 8B (Q6_K)128
Qwen3 8B (Q8_0)104
Gemma 3 12B Instruct · Q4_K_M102
Qwen3 14B · Q4_K_M95.8
Gemma 3 12B Instruct (Q5_K_M)91.0
Gemma 3 12B Instruct (Q8_0)65.5
Mistral Small 3.2 24B Instruct · Q4_K_M62.5
Gemma 3 27B Instruct · Q4_K_M49.4
DECODE TOK/S, SINGLE STREAM
Model
Quant
Weights
Peak VRAM
Headroom
Decode tok/s
Prefill tok/s
TTFT p50
TTFT p95
Works
gpt-oss 20B
Q4_K_M
10.8 GB
11.3 GB
12.7 GB
288.8 ±0.8
12,879.8
19.4 ms
19.9 ms
38% ⚠
Qwen3 30B A3B Instruct 2507
Q4_K_M
17.3 GB
18.0 GB
6.0 GB
259 ±1.2
9,923.8
15.5 ms
16.2 ms
81%
Qwen3 4B Instruct 2507
Q4_K_M
2.3 GB
3.4 GB
20.6 GB
258.4 ±0.4
17,308.2
7.1 ms
7.9 ms
88%
Qwen3 8B (Q3_K_M)
Q3_K_M
3.8 GB
4.7 GB
19.3 GB
186.5 ±0.2
11,108.7
9.3 ms
10.1 ms
63%
GLM 4.7 Flash
Q4_K_M
17.1 GB
17.9 GB
6.1 GB
186.2 ±0.9
7,832.1
18.6 ms
18.8 ms
56%
Llama 3.1 8B Instruct
Q4_K_M
4.6 GB
5.3 GB
18.6 GB
170.6 ±0.2
12,204.5
9.4 ms
9.8 ms
50%
Qwen3 8B
Q4_K_M
4.7 GB
5.5 GB
18.5 GB
163.3 ±0.2
11,988.7
10 ms
10.6 ms
69%
Qwen3 8B (Q5_K_M)
Q5_K_M
5.4 GB
6.1 GB
17.9 GB
144.7 ±0.1
11,655.7
10.7 ms
11.2 ms
50%
Qwen3.5 9B
Q4_K_M
5.3 GB
5.9 GB
18.1 GB
144 ±0.2
9,705.9
56 ms
56.3 ms
69%
Qwen3 8B (Q6_K)
Q6_K
6.3 GB
6.9 GB
17.1 GB
128.3 ±0.1
10,343.4
12.1 ms
12.7 ms
56%
Qwen3 8B (Q8_0)
Q8_0
8.1 GB
8.6 GB
15.4 GB
103.7 ±0.1
12,590.8
13.3 ms
14.2 ms
50%
Gemma 3 12B Instruct
Q4_K_M
6.8 GB
8.9 GB
15.1 GB
102.1 ±0.1
7,276.2
26.7 ms
26.8 ms
94%
Qwen3 14B
Q4_K_M
8.4 GB
9.2 GB
14.8 GB
95.8 ±0.1
6,500.7
16.5 ms
17.1 ms
75%
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
7.9 GB
9.9 GB
14.1 GB
91 ±0.1
7,058.8
28.9 ms
29.2 ms
88%
Gemma 3 12B Instruct (Q8_0)
Q8_0
11.7 GB
13.7 GB
10.3 GB
65.5 ±0
7,822.6
36.5 ms
37.2 ms
88%
Mistral Small 3.2 24B Instruct
Q4_K_M
13.3 GB
14.2 GB
9.8 GB
62.5 ±0
4,129.7
26.1 ms
27.1 ms
81%
Gemma 3 27B Instruct
Q4_K_M
15.4 GB
17.9 GB
6.0 GB
49.4 ±0
3,384.2
49.5 ms
50.6 ms
94%
DOES IT STILL WORK
Accuracy probes after quantisation
Deterministic probes at temperature 0, each with a single correct answer and a programmatic check. Run because a damaged quantisation is faster than an intact one, so the speeds above need an accuracy figure beside them. Looping is scored separately: a model that repeats itself posts a high token rate.
READ THIS BEFORE COMPARING TWO MODELS BY IT
This is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.
Model
Quant
Score
Arithmetic
Degeneration
Factual
Instruction
Json
Language
Logic
Looping
gpt-oss 20B
Q4_K_M
38%
0/3
0/1
3/3
0/3
0/2
2/2
1/2
yes
Llama 3.1 8B Instruct
Q4_K_M
50%
1/3
1/1
3/3
0/3
0/2
2/2
1/2
no
Qwen3 8B (Q5_K_M)
Q5_K_M
50%
0/3
1/1
3/3
0/3
1/2
2/2
1/2
no
Qwen3 8B (Q8_0)
Q8_0
50%
0/3
1/1
3/3
0/3
1/2
2/2
1/2
no
GLM 4.7 Flash
Q4_K_M
56%
1/3
1/1
3/3
0/3
2/2
1/2
1/2
no
Qwen3 8B (Q6_K)
Q6_K
56%
1/3
1/1
3/3
0/3
1/2
2/2
1/2
no
Qwen3 8B (Q3_K_M)
Q3_K_M
63%
1/3
1/1
3/3
0/3
1/2
2/2
2/2
no
Qwen3 8B
Q4_K_M
69%
1/3
1/1
3/3
0/3
2/2
2/2
2/2
no
Qwen3.5 9B
Q4_K_M
69%
2/3
1/1
1/3
3/3
2/2
0/2
2/2
no
Qwen3 14B
Q4_K_M
75%
2/3
1/1
3/3
0/3
2/2
2/2
2/2
no
Qwen3 30B A3B Instruct 2507
Q4_K_M
81%
3/3
1/1
3/3
0/3
2/2
2/2
2/2
no
Mistral Small 3.2 24B Instruct
Q4_K_M
81%
3/3
1/1
3/3
1/3
2/2
1/2
2/2
no
Qwen3 4B Instruct 2507
Q4_K_M
88%
2/3
1/1
3/3
2/3
2/2
2/2
2/2
no
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
88%
2/3
1/1
3/3
3/3
2/2
2/2
1/2
no
Gemma 3 12B Instruct (Q8_0)
Q8_0
88%
2/3
1/1
3/3
3/3
2/2
2/2
1/2
no
Gemma 3 12B Instruct
Q4_K_M
94%
2/3
1/1
3/3
3/3
2/2
2/2
2/2
no
Gemma 3 27B Instruct
Q4_K_M
94%
3/3
1/1
3/3
2/3
2/2
2/2
2/2
no
WHAT THE FAILURES LOOKED LIKE
A bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.
gpt-oss 20B · Q4_K_M — 38%
arith_mul → 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
arith_sub → 1000-387=613. So answer: 613. But let's double-check: 1000-300
arith_mul → What is 47 multiplied by 23? To find the product of 47 and 23,
arith_sub → What is 1000 minus 387? To solve 1000 minus 38
arith_order → Also, what is 12 + 5 * 4? Reply with only the number. What is 1
instr_word → I will give you a list of numbers
instr_count → Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
instr_negation → The number 9 is odd because it
GLM 4.7 Flash · Q4_K_M — 56%
arith_mul → 47 * 23 = ? I need to multiply 47 by 23 and give only the number as the answer
arith_order → 12 + 5 * 4 = 32. 12 + 5 * 4 = 32.
instr_word → BANANA BANANA
instr_count → Red, Green, Blue The user wants exactly three colors. The user specified the format: comma separated, nothing else
instr_negation → YES or NO? YES NO
lang_it → (Non usare parole che contengono la lettera "a"). E
Qwen3 8B (Q6_K) · Q6_K — 56%
arith_sub → What is 1000 minus 387? To solve 1000 minus 38
arith_order → Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
instr_word → If I say "BANANA",
instr_count → Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
instr_negation → The number 9 is odd because it
json_nested → I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDER
Cost of each quantisation step
The same weights at several quantisations, on the same card, in one session. Gives the size of each step in tokens per second and in VRAM, so the rung can be chosen against a measured cost rather than a rule of thumb.
Qwen3 8B
Q3_K_M187
Q4_K_M163
Q5_K_M145
Q6_K128
Q8_0104
DECODE TOK/S
Quant
Weights
Peak VRAM
Decode tok/s
vs fastest rung
Q3_K_M
3.8 GB
4.7 GB
186.5
100%
Q4_K_M
4.7 GB
5.5 GB
163.3
88%
Q5_K_M
5.4 GB
6.1 GB
144.7
78%
Q6_K
6.3 GB
6.9 GB
128.3
69%
Q8_0
8.1 GB
8.6 GB
103.7
56%
Gemma 3 12B Instruct
Q4_K_M102
Q5_K_M91.0
Q8_065.5
DECODE TOK/S
Quant
Weights
Peak VRAM
Decode tok/s
vs fastest rung
Q4_K_M
6.8 GB
8.9 GB
102.1
100%
Q5_K_M
7.9 GB
9.9 GB
91
89%
Q8_0
11.7 GB
13.7 GB
65.5
64%
THE CONTEXT TAX
Decode at depth
Decode measured with the KV cache already filled to the stated depth, not with an empty cache. A missing row means the model failed to load at that context on this GPU.
gpt-oss 20B · Q4_K_M
4,096 ctx272
16,384 ctx246
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
271.9 tok/s
11.1 GB
0.2 GB
5.5 s
16,384
245.6 tok/s
11.4 GB
0.8 GB
5.4 s
Qwen3 30B A3B Instruct 2507 · Q4_K_M
4,096 ctx219
16,384 ctx172
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
218.9 tok/s
18.0 GB
0.4 GB
6.3 s
16,384
172.1 tok/s
19.2 GB
1.5 GB
6.3 s
Qwen3 4B Instruct 2507 · Q4_K_M
4,096 ctx220
16,384 ctx155
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
220 tok/s
3.4 GB
0.6 GB
2.2 s
16,384
154.6 tok/s
5.1 GB
2.3 GB
2.2 s
GLM 4.7 Flash · Q4_K_M
4,096 ctx167
16,384 ctx143
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
167.1 tok/s
17.6 GB
0.2 GB
6.5 s
16,384
143.5 tok/s
18.2 GB
0.8 GB
6.5 s
Llama 3.1 8B Instruct · Q4_K_M
4,096 ctx155
16,384 ctx123
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
155.1 tok/s
5.3 GB
0.5 GB
3.2 s
16,384
122.6 tok/s
6.9 GB
2.0 GB
3.2 s
Qwen3 8B · Q4_K_M
4,096 ctx147
16,384 ctx115
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
147.3 tok/s
5.5 GB
0.6 GB
2.5 s
16,384
114.9 tok/s
7.2 GB
2.3 GB
2.5 s
Qwen3.5 9B · Q4_K_M
4,096 ctx141
16,384 ctx132
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
140.8 tok/s
5.6 GB
0.5 GB
3.1 s
16,384
132.2 tok/s
6.0 GB
2.0 GB
3.1 s
Gemma 3 12B Instruct · Q4_K_M
4,096 ctx95.5
16,384 ctx88.1
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
95.5 tok/s
7.2 GB
1.5 GB
3.1 s
16,384
88.1 tok/s
9.8 GB
6.0 GB
4.1 s
Qwen3 14B · Q4_K_M
4,096 ctx88.8
16,384 ctx74.3
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
88.9 tok/s
9.2 GB
0.6 GB
4.1 s
16,384
74.3 tok/s
11.1 GB
2.5 GB
4.1 s
Mistral Small 3.2 24B Instruct · Q4_K_M
4,096 ctx59.7
16,384 ctx53.0
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
59.7 tok/s
14.2 GB
0.6 GB
4.8 s
16,384
53 tok/s
16.2 GB
2.5 GB
5.5 s
Gemma 3 27B Instruct · Q4_K_M
4,096 ctx47.2
16,384 ctx44.9
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
47.2 tok/s
17.9 GB
1.9 GB
6.1 s
16,384
44.9 tok/s
19.1 GB
7.8 GB
6.1 s
THE RESCUE CURVE
Offload to system RAM
Layers that do not fit stay in system RAM and cross PCIe once per token. Tokens per second at each resident fraction, so a model that overflows the card can be judged on the measured rate rather than on whether it fits.
Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7452 32-Core Processor, 10 cores available to the container, 504 GB RAM, with -t 10 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 24
285.6
9,116.9
11.1 GB
100%
75%
18 / 24
90.1
1,022.2
8.3 GB
32%
50%
12 / 24
55.9
703
5.8 GB
20%
25%
6 / 24
41.1
519.7
3.3 GB
14%
Qwen3 30B A3B Instruct 2507 · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 48
254.7
6,591.8
17.7 GB
100%
75%
36 / 48
86.6
778
13.2 GB
34%
50%
24 / 48
54.7
488.2
9.1 GB
22%
25%
12 / 48
40.1
359.3
4.9 GB
16%
Qwen3 4B Instruct 2507 · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 36
256.7
12,286
2.9 GB
100%
75%
27 / 36
80.1
3,417.4
2.4 GB
31%
50%
18 / 36
52
2,107.3
1.9 GB
20%
25%
9 / 36
35.8
1,524.6
1.4 GB
14%
0%
0 / 36
25
1,204
1.0 GB
10%
GLM 4.7 Flash · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 47
182.9
5,157
17.5 GB
100%
75%
35 / 47
67.4
693.3
13.2 GB
37%
50%
24 / 47
47.9
421.7
9.3 GB
26%
25%
12 / 47
34.4
296.1
5.1 GB
19%
Llama 3.1 8B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 32
169.9
9,614.4
4.8 GB
100%
75%
24 / 32
47
2,179.8
3.8 GB
28%
50%
16 / 32
30.9
1,297.9
2.8 GB
18%
25%
8 / 32
22.2
932.1
1.9 GB
13%
0%
0 / 32
17
721.2
1.0 GB
10%
Qwen3 8B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 36
162.2
9,031.1
5.0 GB
100%
75%
27 / 36
46.4
2,087.1
3.9 GB
29%
50%
18 / 36
28.1
1,250.1
2.9 GB
17%
25%
9 / 36
21.5
882.3
2.0 GB
13%
0%
0 / 36
15.9
692.3
1.1 GB
10%
Qwen3.5 9B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 32
141.6
7,528.1
5.4 GB
100%
75%
24 / 32
38.6
1,806.5
4.4 GB
27%
50%
16 / 32
23.8
1,091.7
3.4 GB
17%
25%
8 / 32
17
796.2
2.4 GB
12%
0%
0 / 32
11.8
624.2
1.6 GB
8%
Gemma 3 12B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 48
101.6
5,851.7
7.4 GB
100%
75%
36 / 48
31.3
1,325.8
5.9 GB
31%
50%
24 / 48
19.3
814
4.5 GB
19%
25%
12 / 48
12.9
582.5
3.0 GB
13%
0%
0 / 48
9.9
453.8
1.6 GB
10%
Qwen3 14B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 40
95.5
5,514.5
8.6 GB
100%
75%
30 / 40
26.6
1,217.7
6.6 GB
28%
50%
20 / 40
16.8
721.1
4.8 GB
18%
25%
10 / 40
12.1
507.6
3.0 GB
13%
0%
0 / 40
8.3
400.4
1.2 GB
9%
Mistral Small 3.2 24B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 40
62.4
3,768.3
13.6 GB
100%
75%
30 / 40
16.4
755.2
10.2 GB
26%
50%
20 / 40
10.8
448.4
7.1 GB
17%
25%
10 / 40
7.7
318.1
4.4 GB
12%
Gemma 3 27B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 62
49.3
2,863.4
16.2 GB
100%
75%
46 / 62
13.8
621.3
12.2 GB
28%
50%
31 / 62
8.9
373.5
8.8 GB
18%
25%
16 / 62
6.2
269
5.4 GB
13%
UNDER LOAD
Concurrency and latency
Aggregate throughput against per-stream rate as concurrent streams increase, with TTFT p95 at each level. Aggregate rises while each stream slows; sizing a deployment needs both columns.
gpt-oss 20B · Q4_K_M — peaks at 862.1 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
264.6 tok/s
264.6 tok/s
31.7 ms
32.2 ms
20
0
2
211.7 tok/s
105.8 tok/s
57.6 ms
59.6 ms
20
0
4
322 tok/s
80.5 tok/s
77.9 ms
96.3 ms
20
0
8
381 tok/s
47.6 tok/s
116.8 ms
139.3 ms
24
0
16
862.1 tok/s
53.9 tok/s
132.6 ms
149.9 ms
48
0
Between 8 and 16 streams the aggregate rose 2.26× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 784.9 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
242 tok/s
242 tok/s
28.1 ms
33.2 ms
20
0
2
399.6 tok/s
199.8 tok/s
31.5 ms
64.1 ms
20
0
4
249.9 tok/s
62.5 tok/s
69 ms
97.4 ms
20
0
8
784.9 tok/s
98.1 tok/s
82.5 ms
99.8 ms
24
0
Between 4 and 8 streams the aggregate rose 3.14× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,321.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
241.1 tok/s
241.1 tok/s
12.9 ms
16.4 ms
20
0
2
197.4 tok/s
98.7 tok/s
22.8 ms
43.8 ms
20
0
4
350.8 tok/s
87.7 tok/s
42.9 ms
99.4 ms
20
0
8
439.9 tok/s
55 tok/s
87.6 ms
92.4 ms
24
0
16
1,321.4 tok/s
82.6 tok/s
137 ms
254 ms
48
0
Between 8 and 16 streams the aggregate rose 3.00× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 1,135.9 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
177.9 tok/s
177.9 tok/s
16.3 ms
17.7 ms
20
0
2
153.4 tok/s
76.7 tok/s
29.9 ms
50.5 ms
20
0
4
267.8 tok/s
67 tok/s
43.8 ms
90.9 ms
20
0
8
331.6 tok/s
41.5 tok/s
118.9 ms
177.9 ms
24
0
16
1,135.9 tok/s
71 tok/s
135.8 ms
161 ms
48
0
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 846.9 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
178.5 tok/s
178.5 tok/s
35.8 ms
36.5 ms
20
0
2
104.1 tok/s
52.1 tok/s
68.3 ms
75.4 ms
20
0
4
192.7 tok/s
48.2 tok/s
82.1 ms
106.5 ms
20
0
8
247.3 tok/s
30.9 tok/s
169.9 ms
215.1 ms
24
0
16
846.9 tok/s
52.9 tok/s
141.6 ms
149.4 ms
48
0
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 1,189.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
164.1 tok/s
164.1 tok/s
14.7 ms
15.8 ms
20
0
2
146.3 tok/s
73.2 tok/s
27.1 ms
45.2 ms
20
0
4
268.4 tok/s
67.1 tok/s
52.6 ms
98.2 ms
20
0
8
344.6 tok/s
43.1 tok/s
82.5 ms
153.9 ms
24
0
16
1,189.4 tok/s
74.3 tok/s
130.8 ms
170.2 ms
48
0
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 1,073.1 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
156.3 tok/s
156.3 tok/s
16.3 ms
17.5 ms
20
0
2
136.5 tok/s
68.2 tok/s
29.5 ms
48.8 ms
20
0
4
247.4 tok/s
61.9 tok/s
42.4 ms
49.3 ms
20
0
8
319 tok/s
39.9 tok/s
98.5 ms
132.1 ms
24
0
16
1,073.1 tok/s
67.1 tok/s
130.9 ms
140.5 ms
48
0
Between 8 and 16 streams the aggregate rose 3.36× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 986 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
139.7 tok/s
139.7 tok/s
16.5 ms
18 ms
20
0
2
123.8 tok/s
61.9 tok/s
31 ms
51.6 ms
20
0
4
228.5 tok/s
57.1 tok/s
58.4 ms
89.5 ms
20
0
8
291.6 tok/s
36.4 tok/s
152.6 ms
155.7 ms
24
0
16
986 tok/s
61.6 tok/s
189.3 ms
264.3 ms
48
0
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 707.3 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
136.8 tok/s
136.8 tok/s
58.3 ms
126.6 ms
20
0
2
116.2 tok/s
58.1 tok/s
127.3 ms
192.6 ms
20
0
4
205.2 tok/s
51.3 tok/s
243.3 ms
500.8 ms
20
0
8
247.9 tok/s
31 tok/s
467.8 ms
736.5 ms
24
0
16
707.3 tok/s
44.2 tok/s
839.4 ms
961.8 ms
48
0
Between 8 and 16 streams the aggregate rose 2.85× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 939.5 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
124.3 tok/s
124.3 tok/s
18.4 ms
20.7 ms
20
0
2
111.6 tok/s
55.8 tok/s
33.5 ms
53.1 ms
20
0
4
208.3 tok/s
52.1 tok/s
69.1 ms
106.1 ms
20
0
8
269.1 tok/s
33.6 tok/s
120 ms
144.1 ms
24
0
16
939.5 tok/s
58.7 tok/s
158 ms
174.3 ms
48
0
Between 8 and 16 streams the aggregate rose 3.49× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 895.7 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
101.3 tok/s
101.3 tok/s
18.8 ms
20.1 ms
20
0
2
92.7 tok/s
46.4 tok/s
35 ms
54.2 ms
20
0
4
174.6 tok/s
43.7 tok/s
45.1 ms
71.4 ms
20
0
8
225.2 tok/s
28.1 tok/s
123 ms
134.6 ms
24
0
16
895.7 tok/s
56 tok/s
133.1 ms
139.8 ms
48
0
Between 8 and 16 streams the aggregate rose 3.98× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 660.9 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
97.9 tok/s
97.9 tok/s
37.7 ms
53.9 ms
20
0
2
85.2 tok/s
42.6 tok/s
70.9 ms
136 ms
20
0
4
156.3 tok/s
39.1 tok/s
100.6 ms
353.8 ms
20
0
8
201.3 tok/s
25.2 tok/s
330.7 ms
560.5 ms
24
0
16
660.9 tok/s
41.3 tok/s
304.4 ms
523.7 ms
48
0
Between 8 and 16 streams the aggregate rose 3.28× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 743.5 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
93.6 tok/s
93.6 tok/s
23.9 ms
25.5 ms
20
0
2
84.8 tok/s
42.4 tok/s
44.7 ms
67.2 ms
20
0
4
159.3 tok/s
39.8 tok/s
65.3 ms
75.2 ms
20
0
8
207.3 tok/s
25.9 tok/s
134.2 ms
181.9 ms
24
0
16
743.5 tok/s
46.5 tok/s
211 ms
236.8 ms
48
0
Between 8 and 16 streams the aggregate rose 3.59× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 639.3 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
88 tok/s
88 tok/s
39.7 ms
54.2 ms
20
0
2
78.5 tok/s
39.2 tok/s
75.2 ms
145.3 ms
20
0
4
143.7 tok/s
35.9 tok/s
167.2 ms
367.1 ms
20
0
8
185.4 tok/s
23.2 tok/s
309.9 ms
506.7 ms
24
0
16
639.3 tok/s
40 tok/s
310 ms
585.2 ms
48
0
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 356.3 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
64.2 tok/s
64.2 tok/s
46.2 ms
65.4 ms
20
0
2
119.9 tok/s
59.9 tok/s
93.1 ms
121.5 ms
20
0
4
110.5 tok/s
27.6 tok/s
173.8 ms
303 ms
20
0
8
356.3 tok/s
44.5 tok/s
295 ms
429.8 ms
24
0
Between 4 and 8 streams the aggregate rose 3.22× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 292.1 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
61.9 tok/s
61.9 tok/s
31.1 ms
39.2 ms
20
0
2
117.1 tok/s
58.5 tok/s
51.4 ms
96.2 ms
20
0
4
111.2 tok/s
27.8 tok/s
124.7 ms
127 ms
20
0
8
292.1 tok/s
36.5 tok/s
205.2 ms
212.9 ms
24
0
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Gemma 3 27B Instruct · Q4_K_M — peaks at 221.7 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
48.6 tok/s
48.6 tok/s
65.6 ms
90.5 ms
20
0
2
90.1 tok/s
45.1 tok/s
94.8 ms
180 ms
20
0
4
84.3 tok/s
21.1 tok/s
327.8 ms
463.9 ms
20
0
8
221.7 tok/s
27.7 tok/s
375.8 ms
449 ms
24
0
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
ONE FLAG
Flash attention, measured at depth
Flash attention on and off, at depth. LM Studio and Ollama enable it by default. The effect depends on the model and on cache size, so it is measured with a filled KV cache rather than an empty one.
gpt-oss 20B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
250.7
4,999.9
11.8 GB
3,968
on
272.1
8,532.1
11.4 GB
3,968
Qwen3 30B A3B Instruct 2507 · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
197.3
3,146.2
18.4 GB
3,968
on
218.9
5,947.8
18.3 GB
3,968
Qwen3 4B Instruct 2507 · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
204.9
4,501.9
3.8 GB
3,968
on
220
9,815.4
3.7 GB
3,968
GLM 4.7 Flash · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
95.2
2,856.5
17.9 GB
3,968
on
167.3
3,970.8
17.9 GB
3,968
Llama 3.1 8B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
148
4,432.8
5.7 GB
3,968
on
155.1
8,314.7
5.5 GB
3,968
Qwen3 8B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
140.2
4,033.9
5.8 GB
3,968
on
147.3
7,762.8
5.7 GB
3,968
Qwen3.5 9B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
138.8
6,382.9
5.9 GB
3,968
on
140.7
7,066.9
5.9 GB
3,968
Gemma 3 12B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
92.9
4,623.6
8.5 GB
3,968
on
95.3
6,008.2
8.5 GB
3,968
Qwen3 14B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
85.3
2,726.9
9.6 GB
3,968
on
88.8
4,645.8
9.4 GB
3,968
Mistral Small 3.2 24B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
58.4
2,424.5
14.5 GB
3,968
on
59.7
3,535.2
14.4 GB
3,968
Gemma 3 27B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
46.4
2,330.7
17.4 GB
3,968
on
47.2
2,973.3
17.3 GB
3,968
WATTS AND MONEY
Power and cost per million tokens
Power sampled at the card during generation, not the TDP from the spec sheet. Cost per million tokens uses the hourly rate this machine was billed at; for owned hardware, apply your own electricity price to the same energy figures.
gpt-oss 20B · Q4_K_M0.664
Qwen3 30B A3B Instruct 2507 · Q4_K_M0.740
Qwen3 4B Instruct 2507 · Q4_K_M0.742
Qwen3 8B (Q3_K_M)1.03
GLM 4.7 Flash · Q4_K_M1.03
Llama 3.1 8B Instruct · Q4_K_M1.12
Qwen3 8B · Q4_K_M1.17
Qwen3 8B (Q5_K_M)1.32
Qwen3.5 9B · Q4_K_M1.33
Qwen3 8B (Q6_K)1.49
Qwen3 8B (Q8_0)1.85
Gemma 3 12B Instruct · Q4_K_M1.88
Qwen3 14B · Q4_K_M2.00
Gemma 3 12B Instruct (Q5_K_M)2.11
Gemma 3 12B Instruct (Q8_0)2.92
Mistral Small 3.2 24B Instruct · Q4_K_M3.07
Gemma 3 27B Instruct · Q4_K_M3.88
$ PER MILLION TOKENS, ONE STREAM — CHEAPER IS SHORTER
Model
Quant
Avg W
Peak W
Peak °C
tok/W
kWh/Mtok
$/Mtok single
$/Mtok batched
gpt-oss 20B
Q4_K_M
210
338
47
1.38
0.202
$0.66
$0.22
Qwen3 30B A3B Instruct 2507
Q4_K_M
181
307
45
1.43
0.195
$0.74
$0.24
Qwen3 4B Instruct 2507
Q4_K_M
268
332
49
0.97
0.288
$0.74
$0.15
Qwen3 8B (Q3_K_M)
Q3_K_M
338
414
55
0.55
0.504
$1.03
$0.17
GLM 4.7 Flash
Q4_K_M
187
302
44
1
0.278
$1.03
$0.23
Llama 3.1 8B Instruct
Q4_K_M
299
361
49
0.57
0.486
$1.12
$0.16
Qwen3 8B
Q4_K_M
298
356
51
0.55
0.506
$1.17
$0.18
Qwen3 8B (Q5_K_M)
Q5_K_M
298
359
52
0.49
0.572
$1.32
$0.19
Qwen3.5 9B
Q4_K_M
291
357
52
0.5
0.562
$1.33
$0.27
Qwen3 8B (Q6_K)
Q6_K
326
390
57
0.39
0.707
$1.49
$0.2
Qwen3 8B (Q8_0)
Q8_0
267
314
49
0.39
0.715
$1.85
$0.21
Gemma 3 12B Instruct
Q4_K_M
302
366
51
0.34
0.823
$1.88
$0.29
Qwen3 14B
Q4_K_M
316
379
53
0.3
0.915
$2
$0.26
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
305
362
53
0.3
0.932
$2.11
$0.3
Gemma 3 12B Instruct (Q8_0)
Q8_0
266
329
53
0.25
1.129
$2.92
$0.54
Mistral Small 3.2 24B Instruct
Q4_K_M
329
413
59
0.19
1.461
$3.07
$0.66
Gemma 3 27B Instruct
Q4_K_M
323
409
59
0.15
1.815
$3.88
$0.86
THE RUN, TAKEN TOGETHER
Relationships across the whole run
Relationships fitted across the whole run rather than measurements of one model: the card's decode constants, the runtime's fixed VRAM overhead, and how TTFT scales with concurrency. Each states its n and its correlation coefficient, and a fit with a weak correlation is quoted without a line drawn through it.
DECODE CONSTANTS: FIXED COST AND BANDWIDTH
At batch one, decode reads every weight once per token. Time per token is therefore a fixed cost plus weight bytes divided by bandwidth, which is a straight line when milliseconds per token is plotted against gigabytes. Fitted over 5 quantisations of Qwen3 8B. r = 1.000 across 5 models.
MILLISECONDS PER TOKEN
WEIGHTS, GB
IMPLIED BANDWIDTH
989 GB/s
inverse of the slope
FIXED PER TOKEN
1.43 ms
intercept; not explained by size
SO A 14 GB MODEL
64.1 tok/s
predicted, not measured
An implied bandwidth above the card’s rated figure does not mean the card is faster than rated. It means the file size on disk is not exactly the number of bytes the runtime moved per token. Treat the two constants as the pair that reproduces these measurements. The 14 GB figure is what they predict; no 14 GB model was measured in this run.
FIXED VRAM OVERHEAD BEFORE THE WEIGHTS
Peak VRAM against weight size fits a line of slope near one. The intercept is everything that is not weights: KV cache, compute buffers and the allocator’s own reservation. A fit estimate that assumes a zero intercept will report that a model fits when it does not. r = 0.991 across 17 models.
Time to first token at the 95th percentile against concurrent streams, on Llama 3.1 8B Instruct · Q4_K_M. Aggregate throughput does not show this: it rises while each individual request waits longer. r = 0.89 across 5 models: a relationship, with visible scatter.
TTFT P95, MS
CONCURRENT STREAMS
EACH ADDED STREAM
+10 ms
onto the 95th percentile
AT ONE STREAM
36 ms
intercept, one stream
THE MACHINE
What this ran on, in full
The full hardware and software configuration these numbers were taken on. Fields that require privileges the harness did not have — the DMI table is not readable inside a container — are marked unavailable rather than omitted, since a missing field and an unreadable one are different.
The NVIDIA GeForce RTX 4090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.
You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.
A pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-4090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →