Everything below is also a PDF — every table, every chart, the machine and the method, in one document you can keep or forward. The raw JSON is the same measurements as data, for anyone who would rather check them than read them.
Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.
Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIX
What runs, and how fast
Decode: single-stream generation, model fully resident on the GPU. Peak VRAM: the maximum the driver reported during the run, not the weight size — the difference is the runtime's own allocation. TTFT p95: time to first token at the 95th percentile, measured under the load stated in the concurrency table.
gpt-oss 20B · Q4_K_M227
Qwen3 30B A3B Instruct 2507 · Q4_K_M208
Qwen3 4B Instruct 2507 · Q4_K_M206
GLM 4.7 Flash · Q4_K_M149
Llama 3.1 8B Instruct · Q4_K_M145
Qwen3 8B · Q4_K_M138
Qwen3 8B (Q5_K_M)124
Qwen3.5 9B · Q4_K_M124
Qwen3 8B (Q3_K_M)119
Qwen3 8B (Q6_K)106
Qwen3 8B (Q8_0)92.9
Gemma 3 12B Instruct · Q4_K_M86.4
Qwen3 14B · Q4_K_M81.5
Gemma 3 12B Instruct (Q5_K_M)77.8
Gemma 3 12B Instruct (Q8_0)58.2
Mistral Small 3.2 24B Instruct · Q4_K_M54.0
Gemma 3 27B Instruct · Q4_K_M42.6
DECODE TOK/S, SINGLE STREAM
Model
Quant
Weights
Peak VRAM
Headroom
Decode tok/s
Prefill tok/s
TTFT p50
TTFT p95
Works
gpt-oss 20B
Q4_K_M
10.8 GB
11.2 GB
12.8 GB
227.4 ±1.1
5,859.2
25.9 ms
28.7 ms
44%
Qwen3 30B A3B Instruct 2507
Q4_K_M
17.3 GB
17.9 GB
6.1 GB
207.6 ±0.8
4,323.8
20.3 ms
22.4 ms
88%
Qwen3 4B Instruct 2507
Q4_K_M
2.3 GB
3.3 GB
20.7 GB
206.4 ±0.9
8,258.2
9.4 ms
14.9 ms
88%
GLM 4.7 Flash
Q4_K_M
17.1 GB
17.5 GB
6.5 GB
149.3 ±0.9
3,653.8
22.9 ms
25.7 ms
56% ⚠
Llama 3.1 8B Instruct
Q4_K_M
4.6 GB
5.2 GB
18.8 GB
145.3 ±0.3
5,189.7
12.8 ms
13.4 ms
50%
Qwen3 8B
Q4_K_M
4.7 GB
5.3 GB
18.7 GB
138 ±0.2
5,001.2
13.3 ms
14 ms
63%
Qwen3 8B (Q5_K_M)
Q5_K_M
5.4 GB
6.0 GB
18.0 GB
123.9 ±0.1
4,856.3
14 ms
14.6 ms
56%
Qwen3.5 9B
Q4_K_M
5.3 GB
5.6 GB
18.4 GB
123.7 ±0.1
4,139.6
64.7 ms
72.8 ms
69%
Qwen3 8B (Q3_K_M)
Q3_K_M
3.8 GB
4.5 GB
19.5 GB
119.1 ±0.1
4,628.2
13.3 ms
14.6 ms
63%
Qwen3 8B (Q6_K)
Q6_K
6.3 GB
6.7 GB
17.3 GB
105.9 ±0.1
4,399.8
15.5 ms
16.2 ms
56%
Qwen3 8B (Q8_0)
Q8_0
8.1 GB
8.5 GB
15.5 GB
92.9 ±0.2
5,214.1
17.2 ms
19.1 ms
50%
Gemma 3 12B Instruct
Q4_K_M
6.8 GB
8.7 GB
15.3 GB
86.4 ±0.1
3,243.4
41.4 ms
42 ms
94%
Qwen3 14B
Q4_K_M
8.4 GB
9.0 GB
15.0 GB
81.5 ±0.1
2,987.5
22.2 ms
23.2 ms
75%
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
7.9 GB
9.8 GB
14.2 GB
77.8 ±0
3,170.8
44.8 ms
45.8 ms
88%
Gemma 3 12B Instruct (Q8_0)
Q8_0
11.7 GB
13.6 GB
10.4 GB
58.2 ±0.1
3,421.2
42.8 ms
44 ms
88%
Mistral Small 3.2 24B Instruct
Q4_K_M
13.3 GB
14.1 GB
9.9 GB
54 ±0
1,915
32.3 ms
33.7 ms
81%
Gemma 3 27B Instruct
Q4_K_M
15.4 GB
17.8 GB
6.2 GB
42.6 ±0
1,482.6
84.7 ms
85.2 ms
94%
DOES IT STILL WORK
Accuracy probes after quantisation
Deterministic probes at temperature 0, each with a single correct answer and a programmatic check. Run because a damaged quantisation is faster than an intact one, so the speeds above need an accuracy figure beside them. Looping is scored separately: a model that repeats itself posts a high token rate.
READ THIS BEFORE COMPARING TWO MODELS BY IT
This is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.
Model
Quant
Score
Arithmetic
Degeneration
Factual
Instruction
Json
Language
Logic
Looping
gpt-oss 20B
Q4_K_M
44%
0/3
1/1
2/3
0/3
0/2
2/2
2/2
no
Llama 3.1 8B Instruct
Q4_K_M
50%
1/3
1/1
3/3
0/3
0/2
2/2
1/2
no
Qwen3 8B (Q8_0)
Q8_0
50%
0/3
1/1
3/3
0/3
1/2
2/2
1/2
no
GLM 4.7 Flash
Q4_K_M
56%
1/3
0/1
3/3
0/3
2/2
2/2
1/2
yes
Qwen3 8B (Q5_K_M)
Q5_K_M
56%
1/3
1/1
3/3
0/3
1/2
2/2
1/2
no
Qwen3 8B (Q6_K)
Q6_K
56%
1/3
1/1
3/3
0/3
1/2
2/2
1/2
no
Qwen3 8B
Q4_K_M
63%
0/3
1/1
3/3
0/3
2/2
2/2
2/2
no
Qwen3 8B (Q3_K_M)
Q3_K_M
63%
1/3
1/1
3/3
0/3
1/2
2/2
2/2
no
Qwen3.5 9B
Q4_K_M
69%
2/3
1/1
1/3
3/3
2/2
0/2
2/2
no
Qwen3 14B
Q4_K_M
75%
2/3
1/1
3/3
0/3
2/2
2/2
2/2
no
Mistral Small 3.2 24B Instruct
Q4_K_M
81%
3/3
1/1
3/3
1/3
2/2
1/2
2/2
no
Qwen3 30B A3B Instruct 2507
Q4_K_M
88%
3/3
1/1
3/3
1/3
2/2
2/2
2/2
no
Qwen3 4B Instruct 2507
Q4_K_M
88%
2/3
1/1
3/3
2/3
2/2
2/2
2/2
no
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
88%
2/3
1/1
3/3
3/3
2/2
2/2
1/2
no
Gemma 3 12B Instruct (Q8_0)
Q8_0
88%
2/3
1/1
3/3
3/3
2/2
2/2
1/2
no
Gemma 3 12B Instruct
Q4_K_M
94%
3/3
1/1
3/3
3/3
2/2
2/2
1/2
no
Gemma 3 27B Instruct
Q4_K_M
94%
3/3
1/1
3/3
2/3
2/2
2/2
2/2
no
WHAT THE FAILURES LOOKED LIKE
A bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.
gpt-oss 20B · Q4_K_M — 44%
arith_mul → 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
arith_sub → 1000-387=613. So answer: 613. But let's double-check: 1000-387
json_nested → I need to make sure that the JSON is valid and properly formatted. Also, ensure that the keys are in double quotes and the boolean value is lowercase. Let me c
Qwen3 8B (Q6_K) · Q6_K — 56%
arith_sub → What is 1000 minus 387? To solve 1000 minus 38
arith_order → Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
instr_word → If I say "BANANA",
instr_count → Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
instr_negation → The number 9 is odd because it
json_nested → I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDER
Cost of each quantisation step
The same weights at several quantisations, on the same card, in one session. Gives the size of each step in tokens per second and in VRAM, so the rung can be chosen against a measured cost rather than a rule of thumb.
Qwen3 8B
Q4_K_M138
Q5_K_M124
Q3_K_M119
Q6_K106
Q8_092.9
DECODE TOK/S
Quant
Weights
Peak VRAM
Decode tok/s
vs fastest rung
Q4_K_M
4.7 GB
5.3 GB
138
100%
Q5_K_M
5.4 GB
6.0 GB
123.9
90%
Q3_K_M
3.8 GB
4.5 GB
119.1
86%
Q6_K
6.3 GB
6.7 GB
105.9
77%
Q8_0
8.1 GB
8.5 GB
92.9
67%
Gemma 3 12B Instruct
Q4_K_M86.4
Q5_K_M77.8
Q8_058.2
DECODE TOK/S
Quant
Weights
Peak VRAM
Decode tok/s
vs fastest rung
Q4_K_M
6.8 GB
8.7 GB
86.4
100%
Q5_K_M
7.9 GB
9.8 GB
77.8
90%
Q8_0
11.7 GB
13.6 GB
58.2
67%
THE CONTEXT TAX
Decode at depth
Decode measured with the KV cache already filled to the stated depth, not with an empty cache. A missing row means the model failed to load at that context on this GPU.
gpt-oss 20B · Q4_K_M
4,096 ctx216
16,384 ctx203
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
216.5 tok/s
11.0 GB
0.2 GB
6.1 s
16,384
203.1 tok/s
11.2 GB
0.8 GB
6 s
Qwen3 30B A3B Instruct 2507 · Q4_K_M
4,096 ctx185
16,384 ctx146
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
185.2 tok/s
17.9 GB
0.4 GB
6.8 s
16,384
146 tok/s
19.0 GB
1.5 GB
7 s
Qwen3 4B Instruct 2507 · Q4_K_M
4,096 ctx178
16,384 ctx131
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
177.7 tok/s
3.3 GB
0.6 GB
2 s
16,384
131.5 tok/s
5.0 GB
2.3 GB
2 s
GLM 4.7 Flash · Q4_K_M
4,096 ctx130
16,384 ctx107
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
130 tok/s
17.5 GB
0.2 GB
7.3 s
16,384
107.1 tok/s
18.1 GB
0.8 GB
7 s
Llama 3.1 8B Instruct · Q4_K_M
4,096 ctx133
16,384 ctx108
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
132.7 tok/s
5.2 GB
0.5 GB
3 s
16,384
107.6 tok/s
6.7 GB
2.0 GB
3 s
Qwen3 8B · Q4_K_M
4,096 ctx125
16,384 ctx100
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
125.5 tok/s
5.3 GB
0.6 GB
3 s
16,384
100.1 tok/s
7.0 GB
2.3 GB
3 s
Qwen3.5 9B · Q4_K_M
4,096 ctx121
16,384 ctx115
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
121.3 tok/s
5.5 GB
0.5 GB
3.3 s
16,384
115.3 tok/s
5.8 GB
2.0 GB
3.3 s
Gemma 3 12B Instruct · Q4_K_M
4,096 ctx81.0
16,384 ctx75.5
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
81.1 tok/s
8.7 GB
1.5 GB
4.8 s
16,384
75.5 tok/s
9.6 GB
6.0 GB
4.2 s
Qwen3 14B · Q4_K_M
4,096 ctx76.7
16,384 ctx65.3
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
76.7 tok/s
9.0 GB
0.6 GB
4.2 s
16,384
65.3 tok/s
10.9 GB
2.5 GB
4.3 s
Mistral Small 3.2 24B Instruct · Q4_K_M
4,096 ctx51.5
16,384 ctx46.4
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
51.6 tok/s
14.1 GB
0.6 GB
4.9 s
16,384
46.4 tok/s
16.0 GB
2.5 GB
4.9 s
Gemma 3 27B Instruct · Q4_K_M
4,096 ctx40.4
16,384 ctx39.0
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
Context
Decode at depth
Peak VRAM
KV cache (theory)
Load time
4,096
40.4 tok/s
17.8 GB
1.9 GB
6.5 s
16,384
39.1 tok/s
19.0 GB
7.8 GB
7.8 s
THE RESCUE CURVE
Offload to system RAM
Layers that do not fit stay in system RAM and cross PCIe once per token. Tokens per second at each resident fraction, so a model that overflows the card can be judged on the measured rate rather than on whether it fits.
Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7H12 64-Core Processor, 27 cores available to the container, 1008 GB RAM, with -t 27 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 24
227.2
4,347.9
11.0 GB
100%
75%
18 / 24
83.8
811.3
8.1 GB
37%
50%
12 / 24
52.5
558.9
5.7 GB
23%
25%
6 / 24
41.4
421.7
3.2 GB
18%
Qwen3 30B A3B Instruct 2507 · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 48
205.5
3,094.2
17.6 GB
100%
75%
36 / 48
56.6
500.5
13.0 GB
28%
50%
24 / 48
34.2
332.1
8.9 GB
17%
25%
12 / 48
24.2
242.5
4.8 GB
12%
Qwen3 4B Instruct 2507 · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 36
203.3
6,439.4
2.8 GB
100%
75%
27 / 36
88.6
2,847.8
2.2 GB
44%
50%
18 / 36
55
1,418.5
1.7 GB
27%
25%
9 / 36
41.7
1,350.7
1.3 GB
21%
0%
0 / 36
30.5
1,081.7
0.8 GB
15%
GLM 4.7 Flash · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 47
146.2
2,498.2
17.4 GB
100%
75%
35 / 47
66.8
600
13.0 GB
46%
50%
24 / 47
42.8
387.3
9.1 GB
29%
25%
12 / 47
29.1
277.7
4.9 GB
20%
Llama 3.1 8B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 32
145.2
4,617.7
4.7 GB
100%
75%
24 / 32
45.7
1,470.8
3.7 GB
32%
50%
16 / 32
28.9
986.7
2.7 GB
20%
25%
8 / 32
19.5
708
1.7 GB
13%
0%
0 / 32
15.7
560.5
0.9 GB
11%
Qwen3 8B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 36
137.8
4,315.5
4.8 GB
100%
75%
27 / 36
47.6
1,690.9
3.8 GB
35%
50%
18 / 36
31.9
1,077.6
2.8 GB
23%
25%
9 / 36
22.5
798.3
1.9 GB
16%
0%
0 / 36
17.5
651.4
1.0 GB
13%
Qwen3.5 9B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 32
122.2
3,534.6
5.3 GB
100%
75%
24 / 32
39.5
1,422.2
4.2 GB
32%
50%
16 / 32
25.2
931.2
3.2 GB
21%
25%
8 / 32
17.7
694.2
2.3 GB
14%
0%
0 / 32
13.1
567
1.4 GB
11%
Gemma 3 12B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 48
86.2
2,849.7
7.4 GB
100%
75%
36 / 48
33.1
1,068.6
5.8 GB
39%
50%
24 / 48
21.6
692.2
4.3 GB
25%
25%
12 / 48
14.8
501.5
2.8 GB
17%
0%
0 / 48
11.2
387.3
1.5 GB
13%
Qwen3 14B · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 40
81.3
2,735
8.5 GB
100%
75%
30 / 40
31.2
984.8
6.4 GB
38%
50%
20 / 40
19.4
630
4.6 GB
24%
25%
10 / 40
14.8
460.4
2.8 GB
18%
0%
0 / 40
11.2
370.4
1.1 GB
14%
Mistral Small 3.2 24B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 40
54
1,804.4
13.5 GB
100%
75%
30 / 40
17.5
633.4
10.0 GB
32%
50%
20 / 40
12.1
395.5
7.0 GB
22%
25%
10 / 40
9.4
289.5
3.9 GB
17%
Gemma 3 27B Instruct · Q4_K_M
DECODE TOK/S
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPU
Layers
Decode tok/s
Prefill tok/s
Peak VRAM
vs fully resident
100%
999 / 62
42.5
1,388.5
16.1 GB
100%
75%
46 / 62
14.3
489.9
12.1 GB
34%
50%
31 / 62
6.8
270.3
8.7 GB
16%
25%
16 / 62
5.8
225.9
5.3 GB
14%
UNDER LOAD
Concurrency and latency
Aggregate throughput against per-stream rate as concurrent streams increase, with TTFT p95 at each level. Aggregate rises while each stream slows; sizing a deployment needs both columns.
gpt-oss 20B · Q4_K_M — peaks at 594 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
210.4 tok/s
210.4 tok/s
49.8 ms
55.2 ms
20
0
2
178.3 tok/s
89.1 tok/s
95.3 ms
98.7 ms
20
0
4
261.1 tok/s
65.3 tok/s
128.9 ms
170.3 ms
20
0
8
290.4 tok/s
36.3 tok/s
230.6 ms
244.8 ms
24
0
16
594 tok/s
37.1 tok/s
258.3 ms
324.7 ms
48
0
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 761.2 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
196.8 tok/s
196.8 tok/s
49.4 ms
53.7 ms
20
0
2
129.7 tok/s
64.9 tok/s
98.1 ms
114.3 ms
20
0
4
226.7 tok/s
56.7 tok/s
123.1 ms
185.8 ms
20
0
8
269.3 tok/s
33.7 tok/s
267 ms
355.9 ms
24
0
16
761.2 tok/s
47.6 tok/s
258.3 ms
306.7 ms
48
0
Between 8 and 16 streams the aggregate rose 2.83× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,108.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
195.6 tok/s
195.6 tok/s
19.4 ms
31.8 ms
20
0
2
164.2 tok/s
82.1 tok/s
34.9 ms
57.4 ms
20
0
4
276.6 tok/s
69.1 tok/s
67.3 ms
108.7 ms
20
0
8
327.6 tok/s
41 tok/s
182.9 ms
191.1 ms
24
0
16
1,108.4 tok/s
69.3 tok/s
253.9 ms
398.6 ms
48
0
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 566.2 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
144 tok/s
144 tok/s
63.9 ms
64.8 ms
20
0
2
100.8 tok/s
50.4 tok/s
126.3 ms
134.4 ms
20
0
4
175.1 tok/s
43.8 tok/s
155.7 ms
179.5 ms
20
0
8
208.3 tok/s
26 tok/s
304.1 ms
377.2 ms
24
0
16
566.2 tok/s
35.4 tok/s
326 ms
371.1 ms
48
0
Between 8 and 16 streams the aggregate rose 2.72× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 923.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
140.6 tok/s
140.6 tok/s
26.7 ms
28.4 ms
20
0
2
126.3 tok/s
63.2 tok/s
47.6 ms
65.9 ms
20
0
4
211.9 tok/s
53 tok/s
95.8 ms
118.5 ms
20
0
8
248 tok/s
31 tok/s
173.9 ms
182.2 ms
24
0
16
923.4 tok/s
57.7 tok/s
307 ms
410.3 ms
48
0
Between 8 and 16 streams the aggregate rose 3.72× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 853.5 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
133.1 tok/s
133.1 tok/s
28.8 ms
30.7 ms
20
0
2
117.2 tok/s
58.6 tok/s
51 ms
76.6 ms
20
0
4
197.4 tok/s
49.4 tok/s
84.9 ms
110.3 ms
20
0
8
231.9 tok/s
29 tok/s
206.7 ms
214.1 ms
24
0
16
853.5 tok/s
53.3 tok/s
300.1 ms
491.5 ms
48
0
Between 8 and 16 streams the aggregate rose 3.68× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 836 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
120.1 tok/s
120.1 tok/s
29.3 ms
31 ms
20
0
2
106.6 tok/s
53.3 tok/s
53.7 ms
84.9 ms
20
0
4
182.7 tok/s
45.7 tok/s
84.3 ms
161.7 ms
20
0
8
215.8 tok/s
27 tok/s
225.9 ms
265.3 ms
24
0
16
836 tok/s
52.3 tok/s
393.8 ms
588.2 ms
48
0
Between 8 and 16 streams the aggregate rose 3.87× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 523.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
119.2 tok/s
119.2 tok/s
73.9 ms
158.3 ms
20
0
2
102.8 tok/s
51.4 tok/s
161.8 ms
244.7 ms
20
0
4
165.9 tok/s
41.5 tok/s
380.4 ms
608.2 ms
20
0
8
194.3 tok/s
24.3 tok/s
817.2 ms
971.5 ms
24
0
16
523.4 tok/s
32.7 tok/s
1,327.4 ms
2,100.9 ms
48
0
Between 8 and 16 streams the aggregate rose 2.69× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 895 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
116.2 tok/s
116.2 tok/s
29.5 ms
31.9 ms
20
0
2
104.9 tok/s
52.5 tok/s
54.3 ms
77.5 ms
20
0
4
176.1 tok/s
44 tok/s
86.5 ms
133.5 ms
20
0
8
211.2 tok/s
26.4 tok/s
179.4 ms
242.5 ms
24
0
16
895 tok/s
55.9 tok/s
293.5 ms
334 ms
48
0
Between 8 and 16 streams the aggregate rose 4.24× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 769.8 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
103.6 tok/s
103.6 tok/s
31.6 ms
33.1 ms
20
0
2
93.9 tok/s
46.9 tok/s
59.2 ms
87.3 ms
20
0
4
169.1 tok/s
42.3 tok/s
108.5 ms
131.5 ms
20
0
8
206.4 tok/s
25.8 tok/s
221 ms
238.7 ms
24
0
16
769.8 tok/s
48.1 tok/s
325.6 ms
364.6 ms
48
0
Between 8 and 16 streams the aggregate rose 3.73× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 713.2 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
90.8 tok/s
90.8 tok/s
26.9 ms
28.8 ms
20
0
2
82.7 tok/s
41.4 tok/s
52.5 ms
76 ms
20
0
4
155 tok/s
38.8 tok/s
81.9 ms
166.3 ms
20
0
8
199.2 tok/s
24.9 tok/s
234.4 ms
296.4 ms
24
0
16
713.2 tok/s
44.6 tok/s
313.6 ms
372.8 ms
48
0
Between 8 and 16 streams the aggregate rose 3.58× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 556.4 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
84 tok/s
84 tok/s
59.3 ms
79.7 ms
20
0
2
74 tok/s
37 tok/s
122.7 ms
216.7 ms
20
0
4
122.4 tok/s
30.6 tok/s
179.1 ms
415.3 ms
20
0
8
144.8 tok/s
18.1 tok/s
492.5 ms
669.1 ms
24
0
16
556.4 tok/s
34.8 tok/s
758.6 ms
876 ms
48
0
Between 8 and 16 streams the aggregate rose 3.84× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 613.3 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
79.7 tok/s
79.7 tok/s
43 ms
45.4 ms
20
0
2
72.9 tok/s
36.5 tok/s
90.5 ms
105.9 ms
20
0
4
121.8 tok/s
30.5 tok/s
127.7 ms
187.6 ms
20
0
8
143.7 tok/s
18 tok/s
279.7 ms
296.6 ms
24
0
16
613.3 tok/s
38.3 tok/s
410 ms
431.4 ms
48
0
Between 8 and 16 streams the aggregate rose 4.27× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 517.5 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
75.4 tok/s
75.4 tok/s
65.1 ms
81 ms
20
0
2
67.5 tok/s
33.7 tok/s
131.2 ms
229.9 ms
20
0
4
111.6 tok/s
27.9 tok/s
293.2 ms
399 ms
20
0
8
133 tok/s
16.6 tok/s
632.3 ms
819.6 ms
24
0
16
517.5 tok/s
32.3 tok/s
1,032.8 ms
1,153.7 ms
48
0
Between 8 and 16 streams the aggregate rose 3.89× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 290.2 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
57 tok/s
57 tok/s
60 ms
90.3 ms
20
0
2
107.7 tok/s
53.8 tok/s
102.5 ms
150.8 ms
20
0
4
97.8 tok/s
24.4 tok/s
315.2 ms
379.2 ms
20
0
8
290.2 tok/s
36.3 tok/s
388 ms
401.5 ms
24
0
Between 4 and 8 streams the aggregate rose 2.97× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 437.1 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
53.6 tok/s
53.6 tok/s
61.2 ms
64.1 ms
20
0
2
50.6 tok/s
25.3 tok/s
116.2 ms
142.8 ms
20
0
4
81.8 tok/s
20.4 tok/s
222.8 ms
254 ms
20
0
8
95.1 tok/s
11.9 tok/s
454.8 ms
498.9 ms
24
0
16
437.1 tok/s
27.3 tok/s
764.7 ms
883.2 ms
48
0
Between 8 and 16 streams the aggregate rose 4.60× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 27B Instruct · Q4_K_M — peaks at 110.1 tok/s aggregate
AGGREGATE TOK/S
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
CONCURRENT STREAMS
Streams
Aggregate
Per stream
TTFT p50
TTFT p95
Samples
Failed
1
41.9 tok/s
41.9 tok/s
112.9 ms
143.3 ms
20
0
2
72.5 tok/s
36.3 tok/s
266.2 ms
316 ms
20
0
4
62.9 tok/s
15.7 tok/s
417.8 ms
610 ms
20
0
8
110.1 tok/s
13.8 tok/s
683.8 ms
773.5 ms
24
0
Concurrency was capped by VRAM on this card, not by compute.
ONE FLAG
Flash attention, measured at depth
Flash attention on and off, at depth. LM Studio and Ollama enable it by default. The effect depends on the model and on cache size, so it is measured with a filled KV cache rather than an empty one.
gpt-oss 20B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
191.5
2,748.1
11.7 GB
3,968
on
218.2
4,199.9
11.3 GB
3,968
Qwen3 30B A3B Instruct 2507 · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
141.8
1,749.5
18.2 GB
3,968
on
186.1
2,799.2
18.1 GB
3,968
Qwen3 4B Instruct 2507 · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
151.8
2,756.7
3.6 GB
3,968
on
177.6
5,283.2
3.5 GB
3,968
GLM 4.7 Flash · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
58
1,499.8
17.8 GB
3,968
on
131.1
1,930
17.7 GB
3,968
Llama 3.1 8B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
119.3
2,548
5.6 GB
3,968
on
133.1
4,101.3
5.4 GB
3,968
Qwen3 8B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
111.3
2,243
5.7 GB
3,968
on
125.5
3,668
5.5 GB
3,968
Qwen3.5 9B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
118.2
3,211.4
5.8 GB
3,968
on
121.6
3,458.6
5.7 GB
3,968
Gemma 3 12B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
78
2,330
8.7 GB
3,968
on
81.3
2,764.8
8.3 GB
3,968
Qwen3 14B · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
68.3
1,570.2
9.5 GB
3,968
on
77
2,254.8
9.2 GB
3,968
Mistral Small 3.2 24B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
48.8
1,303.6
14.4 GB
3,968
on
51.9
1,703.5
14.2 GB
3,968
Gemma 3 27B Instruct · Q4_K_M
-fa
Decode tok/s
Prefill tok/s
Peak VRAM
Depth
off
39.6
1,136.3
17.3 GB
3,968
on
40.9
1,401.7
17.2 GB
3,968
WATTS AND MONEY
Power and cost per million tokens
Power sampled at the card during generation, not the TDP from the spec sheet. Cost per million tokens uses the hourly rate this machine was billed at; for owned hardware, apply your own electricity price to the same energy figures.
gpt-oss 20B · Q4_K_M0.611
Qwen3 30B A3B Instruct 2507 · Q4_K_M0.669
Qwen3 4B Instruct 2507 · Q4_K_M0.673
GLM 4.7 Flash · Q4_K_M0.930
Llama 3.1 8B Instruct · Q4_K_M0.956
Qwen3 8B · Q4_K_M1.01
Qwen3 8B (Q5_K_M)1.12
Qwen3.5 9B · Q4_K_M1.12
Qwen3 8B (Q3_K_M)1.17
Qwen3 8B (Q6_K)1.31
Qwen3 8B (Q8_0)1.49
Gemma 3 12B Instruct · Q4_K_M1.61
Qwen3 14B · Q4_K_M1.70
Gemma 3 12B Instruct (Q5_K_M)1.79
Gemma 3 12B Instruct (Q8_0)2.39
Mistral Small 3.2 24B Instruct · Q4_K_M2.57
Gemma 3 27B Instruct · Q4_K_M3.26
$ PER MILLION TOKENS, ONE STREAM — CHEAPER IS SHORTER
Model
Quant
Avg W
Peak W
Peak °C
tok/W
kWh/Mtok
$/Mtok single
$/Mtok batched
gpt-oss 20B
Q4_K_M
248
348
57
0.92
0.302
$0.61
$0.23
Qwen3 30B A3B Instruct 2507
Q4_K_M
237
349
57
0.88
0.317
$0.67
$0.18
Qwen3 4B Instruct 2507
Q4_K_M
300
347
59
0.69
0.404
$0.67
$0.13
GLM 4.7 Flash
Q4_K_M
261
349
58
0.57
0.486
$0.93
$0.25
Llama 3.1 8B Instruct
Q4_K_M
297
348
57
0.49
0.568
$0.96
$0.15
Qwen3 8B
Q4_K_M
305
350
60
0.45
0.613
$1.01
$0.16
Qwen3 8B (Q5_K_M)
Q5_K_M
304
348
62
0.41
0.682
$1.12
$0.17
Qwen3.5 9B
Q4_K_M
302
350
61
0.41
0.678
$1.12
$0.27
Qwen3 8B (Q3_K_M)
Q3_K_M
311
348
63
0.38
0.725
$1.17
$0.16
Qwen3 8B (Q6_K)
Q6_K
306
349
62
0.35
0.803
$1.31
$0.18
Qwen3 8B (Q8_0)
Q8_0
307
350
62
0.3
0.918
$1.49
$0.19
Gemma 3 12B Instruct
Q4_K_M
311
349
62
0.28
0.999
$1.61
$0.25
Qwen3 14B
Q4_K_M
310
348
62
0.26
1.055
$1.7
$0.23
Gemma 3 12B Instruct (Q5_K_M)
Q5_K_M
312
348
63
0.25
1.115
$1.79
$0.27
Gemma 3 12B Instruct (Q8_0)
Q8_0
312
350
63
0.19
1.491
$2.39
$0.48
Mistral Small 3.2 24B Instruct
Q4_K_M
312
349
63
0.17
1.605
$2.57
$0.32
Gemma 3 27B Instruct
Q4_K_M
316
349
64
0.14
2.059
$3.26
$1.26
THE RUN, TAKEN TOGETHER
Relationships across the whole run
Relationships fitted across the whole run rather than measurements of one model: the card's decode constants, the runtime's fixed VRAM overhead, and how TTFT scales with concurrency. Each states its n and its correlation coefficient, and a fit with a weak correlation is quoted without a line drawn through it.
DECODE CONSTANTS: FIXED COST AND BANDWIDTH
At batch one, decode reads every weight once per token. Time per token is therefore a fixed cost plus weight bytes divided by bandwidth, which is a straight line when milliseconds per token is plotted against gigabytes. Fitted over 5 quantisations of Qwen3 8B. r = 0.86 across 5 models: a relationship, with visible scatter.
MILLISECONDS PER TOKEN
WEIGHTS, GB
IMPLIED BANDWIDTH
1,407 GB/s
inverse of the slope
FIXED PER TOKEN
4.75 ms
intercept; not explained by size
SO A 14 GB MODEL
68 tok/s
predicted, not measured
An implied bandwidth above the card’s rated figure does not mean the card is faster than rated. It means the file size on disk is not exactly the number of bytes the runtime moved per token. Treat the two constants as the pair that reproduces these measurements. The 14 GB figure is what they predict; no 14 GB model was measured in this run.
FIXED VRAM OVERHEAD BEFORE THE WEIGHTS
Peak VRAM against weight size fits a line of slope near one. The intercept is everything that is not weights: KV cache, compute buffers and the allocator’s own reservation. A fit estimate that assumes a zero intercept will report that a model fits when it does not. r = 0.991 across 17 models.
Time to first token at the 95th percentile against concurrent streams, on Llama 3.1 8B Instruct · Q4_K_M. Aggregate throughput does not show this: it rises while each individual request waits longer. r = 0.995 across 5 models.
TTFT P95, MS
CONCURRENT STREAMS
EACH ADDED STREAM
+25 ms
onto the 95th percentile
AT ONE STREAM
8 ms
intercept, one stream
THE MACHINE
What this ran on, in full
The full hardware and software configuration these numbers were taken on. Fields that require privileges the harness did not have — the DMI table is not readable inside a container — are marked unavailable rather than omitted, since a missing field and an unreadable one are different.
The NVIDIA GeForce RTX 3090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.
You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.
A pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-3090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →