FitMyLLM
← all rig reports
▸ TAKE THE REPORT WITH YOU

Everything below is also a PDF — every table, every chart, the machine and the method, in one document you can keep or forward. The raw JSON is the same measurements as data, for anyone who would rather check them than read them.

↓ RAW JSON
▸ MEASURED

NVIDIA GeForce RTX 4090

MODELS MEASURED
17
DATA POINTS
549
HOURS ON THE BENCH
2.3
MEASURED ON
2026-08-01
PROVENANCE
24.0 GB VRAMdriver 580.159.04llama.cpp b10156provider runpodcontext 4,096

Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.

Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIX

What runs, and how fast

Decode: single-stream generation, model fully resident on the GPU. Peak VRAM: the maximum the driver reported during the run, not the weight size — the difference is the runtime's own allocation. TTFT p95: time to first token at the 95th percentile, measured under the load stated in the concurrency table.

gpt-oss 20B · Q4_K_M289
Qwen3 30B A3B Instruct 2507 · Q4_K_M259
Qwen3 4B Instruct 2507 · Q4_K_M258
Qwen3 8B (Q3_K_M)187
GLM 4.7 Flash · Q4_K_M186
Llama 3.1 8B Instruct · Q4_K_M171
Qwen3 8B · Q4_K_M163
Qwen3 8B (Q5_K_M)145
Qwen3.5 9B · Q4_K_M144
Qwen3 8B (Q6_K)128
Qwen3 8B (Q8_0)104
Gemma 3 12B Instruct · Q4_K_M102
Qwen3 14B · Q4_K_M95.8
Gemma 3 12B Instruct (Q5_K_M)91.0
Gemma 3 12B Instruct (Q8_0)65.5
Mistral Small 3.2 24B Instruct · Q4_K_M62.5
Gemma 3 27B Instruct · Q4_K_M49.4
DECODE TOK/S, SINGLE STREAM
ModelQuantWeightsPeak VRAMHeadroomDecode tok/sPrefill tok/sTTFT p50TTFT p95Works
gpt-oss 20BQ4_K_M10.8 GB11.3 GB12.7 GB288.8 ±0.812,879.819.4 ms19.9 ms38%
Qwen3 30B A3B Instruct 2507Q4_K_M17.3 GB18.0 GB6.0 GB259 ±1.29,923.815.5 ms16.2 ms81%
Qwen3 4B Instruct 2507Q4_K_M2.3 GB3.4 GB20.6 GB258.4 ±0.417,308.27.1 ms7.9 ms88%
Qwen3 8B (Q3_K_M)Q3_K_M3.8 GB4.7 GB19.3 GB186.5 ±0.211,108.79.3 ms10.1 ms63%
GLM 4.7 FlashQ4_K_M17.1 GB17.9 GB6.1 GB186.2 ±0.97,832.118.6 ms18.8 ms56%
Llama 3.1 8B InstructQ4_K_M4.6 GB5.3 GB18.6 GB170.6 ±0.212,204.59.4 ms9.8 ms50%
Qwen3 8BQ4_K_M4.7 GB5.5 GB18.5 GB163.3 ±0.211,988.710 ms10.6 ms69%
Qwen3 8B (Q5_K_M)Q5_K_M5.4 GB6.1 GB17.9 GB144.7 ±0.111,655.710.7 ms11.2 ms50%
Qwen3.5 9BQ4_K_M5.3 GB5.9 GB18.1 GB144 ±0.29,705.956 ms56.3 ms69%
Qwen3 8B (Q6_K)Q6_K6.3 GB6.9 GB17.1 GB128.3 ±0.110,343.412.1 ms12.7 ms56%
Qwen3 8B (Q8_0)Q8_08.1 GB8.6 GB15.4 GB103.7 ±0.112,590.813.3 ms14.2 ms50%
Gemma 3 12B InstructQ4_K_M6.8 GB8.9 GB15.1 GB102.1 ±0.17,276.226.7 ms26.8 ms94%
Qwen3 14BQ4_K_M8.4 GB9.2 GB14.8 GB95.8 ±0.16,500.716.5 ms17.1 ms75%
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M7.9 GB9.9 GB14.1 GB91 ±0.17,058.828.9 ms29.2 ms88%
Gemma 3 12B Instruct (Q8_0)Q8_011.7 GB13.7 GB10.3 GB65.5 ±07,822.636.5 ms37.2 ms88%
Mistral Small 3.2 24B InstructQ4_K_M13.3 GB14.2 GB9.8 GB62.5 ±04,129.726.1 ms27.1 ms81%
Gemma 3 27B InstructQ4_K_M15.4 GB17.9 GB6.0 GB49.4 ±03,384.249.5 ms50.6 ms94%
DOES IT STILL WORK

Accuracy probes after quantisation

Deterministic probes at temperature 0, each with a single correct answer and a programmatic check. Run because a damaged quantisation is faster than an intact one, so the speeds above need an accuracy figure beside them. Looping is scored separately: a model that repeats itself posts a high token rate.

READ THIS BEFORE COMPARING TWO MODELS BY IT

This is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.

ModelQuantScoreArithmeticDegenerationFactualInstructionJsonLanguageLogicLooping
gpt-oss 20BQ4_K_M38%0/30/13/30/30/22/21/2yes
Llama 3.1 8B InstructQ4_K_M50%1/31/13/30/30/22/21/2no
Qwen3 8B (Q5_K_M)Q5_K_M50%0/31/13/30/31/22/21/2no
Qwen3 8B (Q8_0)Q8_050%0/31/13/30/31/22/21/2no
GLM 4.7 FlashQ4_K_M56%1/31/13/30/32/21/21/2no
Qwen3 8B (Q6_K)Q6_K56%1/31/13/30/31/22/21/2no
Qwen3 8B (Q3_K_M)Q3_K_M63%1/31/13/30/31/22/22/2no
Qwen3 8BQ4_K_M69%1/31/13/30/32/22/22/2no
Qwen3.5 9BQ4_K_M69%2/31/11/33/32/20/22/2no
Qwen3 14BQ4_K_M75%2/31/13/30/32/22/22/2no
Qwen3 30B A3B Instruct 2507Q4_K_M81%3/31/13/30/32/22/22/2no
Mistral Small 3.2 24B InstructQ4_K_M81%3/31/13/31/32/21/22/2no
Qwen3 4B Instruct 2507Q4_K_M88%2/31/13/32/32/22/22/2no
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M88%2/31/13/33/32/22/21/2no
Gemma 3 12B Instruct (Q8_0)Q8_088%2/31/13/33/32/22/21/2no
Gemma 3 12B InstructQ4_K_M94%2/31/13/33/32/22/22/2no
Gemma 3 27B InstructQ4_K_M94%3/31/13/32/32/22/22/2no
WHAT THE FAILURES LOOKED LIKE

A bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.

gpt-oss 20B · Q4_K_M — 38%
  • arith_mul 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
  • arith_sub 1000-387=613. So answer: 613. But let's double-check: 1000-300
  • arith_order 12 + 5 * 4 = 12 + 20 = 32. So answer: 32.
  • instr_word So final answer: "BANANA".
  • instr_count So we need to produce exactly three colours. So we need to decide which three colours to output. We can choose "
  • instr_negation NO The number 9 is
Llama 3.1 8B Instruct · Q4_K_M — 50%
  • arith_mul 1071 #math #multiplication #answer The final answer is: 1071 What is
  • arith_order 61 ## Step 1: First, we need to follow the order of operations, which is Parentheses,
  • instr_word I'll respond with a word that
  • instr_count Red, Blue, Yellow. Answer: Red, Blue, Yellow. This is a list of colours. The
  • instr_negation NO Answer with the single word YES
  • json_person {"name": "Ada", "age": 36} ```python import json def person_info(name, age): return json.dumps({"name": name, "age": age}) print(person_info("Ada", 36)) ``` ```
Qwen3 8B (Q5_K_M) · Q5_K_M — 50%
  • arith_mul To solve 47 multiplied by 23, we can use the standard multiplication algorithm. First, we multiply
  • arith_sub What is 1000 minus 387? Reply with only the number. Okay, let's see
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word If you reply with anything else, you
  • instr_count 1. Red, 2. Blue, 3. Green. 1. Red, 2. Blue
  • instr_negation The number 9 is odd because it
Qwen3 8B (Q8_0) · Q8_0 — 50%
  • arith_mul What is 47 multiplied by 23? To find the product of 47 and 23,
  • arith_sub What is 1000 minus 387? To solve 1000 minus 38
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word I will give you a list of numbers
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
GLM 4.7 Flash · Q4_K_M — 56%
  • arith_mul 47 * 23 = ? I need to multiply 47 by 23 and give only the number as the answer
  • arith_order 12 + 5 * 4 = 32. 12 + 5 * 4 = 32.
  • instr_word BANANA BANANA
  • instr_count Red, Green, Blue The user wants exactly three colors. The user specified the format: comma separated, nothing else
  • instr_negation YES or NO? YES NO
  • lang_it (Non usare parole che contengono la lettera "a"). E
Qwen3 8B (Q6_K) · Q6_K — 56%
  • arith_sub What is 1000 minus 387? To solve 1000 minus 38
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
  • instr_word If I say "BANANA",
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
  • json_nested I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDER

Cost of each quantisation step

The same weights at several quantisations, on the same card, in one session. Gives the size of each step in tokens per second and in VRAM, so the rung can be chosen against a measured cost rather than a rule of thumb.

Qwen3 8B
Q3_K_M187
Q4_K_M163
Q5_K_M145
Q6_K128
Q8_0104
DECODE TOK/S
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q3_K_M3.8 GB4.7 GB186.5100%
Q4_K_M4.7 GB5.5 GB163.388%
Q5_K_M5.4 GB6.1 GB144.778%
Q6_K6.3 GB6.9 GB128.369%
Q8_08.1 GB8.6 GB103.756%
Gemma 3 12B Instruct
Q4_K_M102
Q5_K_M91.0
Q8_065.5
DECODE TOK/S
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q4_K_M6.8 GB8.9 GB102.1100%
Q5_K_M7.9 GB9.9 GB9189%
Q8_011.7 GB13.7 GB65.564%
THE CONTEXT TAX

Decode at depth

Decode measured with the KV cache already filled to the stated depth, not with an empty cache. A missing row means the model failed to load at that context on this GPU.

gpt-oss 20B · Q4_K_M
4,096 ctx272
16,384 ctx246
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096271.9 tok/s11.1 GB0.2 GB5.5 s
16,384245.6 tok/s11.4 GB0.8 GB5.4 s
Qwen3 30B A3B Instruct 2507 · Q4_K_M
4,096 ctx219
16,384 ctx172
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096218.9 tok/s18.0 GB0.4 GB6.3 s
16,384172.1 tok/s19.2 GB1.5 GB6.3 s
Qwen3 4B Instruct 2507 · Q4_K_M
4,096 ctx220
16,384 ctx155
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096220 tok/s3.4 GB0.6 GB2.2 s
16,384154.6 tok/s5.1 GB2.3 GB2.2 s
GLM 4.7 Flash · Q4_K_M
4,096 ctx167
16,384 ctx143
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096167.1 tok/s17.6 GB0.2 GB6.5 s
16,384143.5 tok/s18.2 GB0.8 GB6.5 s
Llama 3.1 8B Instruct · Q4_K_M
4,096 ctx155
16,384 ctx123
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096155.1 tok/s5.3 GB0.5 GB3.2 s
16,384122.6 tok/s6.9 GB2.0 GB3.2 s
Qwen3 8B · Q4_K_M
4,096 ctx147
16,384 ctx115
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096147.3 tok/s5.5 GB0.6 GB2.5 s
16,384114.9 tok/s7.2 GB2.3 GB2.5 s
Qwen3.5 9B · Q4_K_M
4,096 ctx141
16,384 ctx132
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096140.8 tok/s5.6 GB0.5 GB3.1 s
16,384132.2 tok/s6.0 GB2.0 GB3.1 s
Gemma 3 12B Instruct · Q4_K_M
4,096 ctx95.5
16,384 ctx88.1
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09695.5 tok/s7.2 GB1.5 GB3.1 s
16,38488.1 tok/s9.8 GB6.0 GB4.1 s
Qwen3 14B · Q4_K_M
4,096 ctx88.8
16,384 ctx74.3
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09688.9 tok/s9.2 GB0.6 GB4.1 s
16,38474.3 tok/s11.1 GB2.5 GB4.1 s
Mistral Small 3.2 24B Instruct · Q4_K_M
4,096 ctx59.7
16,384 ctx53.0
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09659.7 tok/s14.2 GB0.6 GB4.8 s
16,38453 tok/s16.2 GB2.5 GB5.5 s
Gemma 3 27B Instruct · Q4_K_M
4,096 ctx47.2
16,384 ctx44.9
DECODE TOK/S, CACHE ALREADY FULL TO THAT DEPTH
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09647.2 tok/s17.9 GB1.9 GB6.1 s
16,38444.9 tok/s19.1 GB7.8 GB6.1 s
THE RESCUE CURVE

Offload to system RAM

Layers that do not fit stay in system RAM and cross PCIe once per token. Tokens per second at each resident fraction, so a model that overflows the card can be judged on the measured rate rather than on whether it fits.

Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7452 32-Core Processor, 10 cores available to the container, 504 GB RAM, with -t 10 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
DECODE TOK/S
014328625%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 24285.69,116.911.1 GB100%
75%18 / 2490.11,022.28.3 GB32%
50%12 / 2455.97035.8 GB20%
25%6 / 2441.1519.73.3 GB14%
Qwen3 30B A3B Instruct 2507 · Q4_K_M
DECODE TOK/S
012725525%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 48254.76,591.817.7 GB100%
75%36 / 4886.677813.2 GB34%
50%24 / 4854.7488.29.1 GB22%
25%12 / 4840.1359.34.9 GB16%
Qwen3 4B Instruct 2507 · Q4_K_M
DECODE TOK/S
01282570%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36256.712,2862.9 GB100%
75%27 / 3680.13,417.42.4 GB31%
50%18 / 36522,107.31.9 GB20%
25%9 / 3635.81,524.61.4 GB14%
0%0 / 36251,2041.0 GB10%
GLM 4.7 Flash · Q4_K_M
DECODE TOK/S
091.518325%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 47182.95,15717.5 GB100%
75%35 / 4767.4693.313.2 GB37%
50%24 / 4747.9421.79.3 GB26%
25%12 / 4734.4296.15.1 GB19%
Llama 3.1 8B Instruct · Q4_K_M
DECODE TOK/S
085.01700%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32169.99,614.44.8 GB100%
75%24 / 32472,179.83.8 GB28%
50%16 / 3230.91,297.92.8 GB18%
25%8 / 3222.2932.11.9 GB13%
0%0 / 3217721.21.0 GB10%
Qwen3 8B · Q4_K_M
DECODE TOK/S
081.11620%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36162.29,031.15.0 GB100%
75%27 / 3646.42,087.13.9 GB29%
50%18 / 3628.11,250.12.9 GB17%
25%9 / 3621.5882.32.0 GB13%
0%0 / 3615.9692.31.1 GB10%
Qwen3.5 9B · Q4_K_M
DECODE TOK/S
070.81420%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32141.67,528.15.4 GB100%
75%24 / 3238.61,806.54.4 GB27%
50%16 / 3223.81,091.73.4 GB17%
25%8 / 3217796.22.4 GB12%
0%0 / 3211.8624.21.6 GB8%
Gemma 3 12B Instruct · Q4_K_M
DECODE TOK/S
050.81020%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 48101.65,851.77.4 GB100%
75%36 / 4831.31,325.85.9 GB31%
50%24 / 4819.38144.5 GB19%
25%12 / 4812.9582.53.0 GB13%
0%0 / 489.9453.81.6 GB10%
Qwen3 14B · Q4_K_M
DECODE TOK/S
047.895.50%25%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4095.55,514.58.6 GB100%
75%30 / 4026.61,217.76.6 GB28%
50%20 / 4016.8721.14.8 GB18%
25%10 / 4012.1507.63.0 GB13%
0%0 / 408.3400.41.2 GB9%
Mistral Small 3.2 24B Instruct · Q4_K_M
DECODE TOK/S
031.262.425%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4062.43,768.313.6 GB100%
75%30 / 4016.4755.210.2 GB26%
50%20 / 4010.8448.47.1 GB17%
25%10 / 407.7318.14.4 GB12%
Gemma 3 27B Instruct · Q4_K_M
DECODE TOK/S
024.649.325%50%75%100%
SHARE OF THE MODEL RESIDENT ON THE GPU
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 6249.32,863.416.2 GB100%
75%46 / 6213.8621.312.2 GB28%
50%31 / 628.9373.58.8 GB18%
25%16 / 626.22695.4 GB13%
UNDER LOAD

Concurrency and latency

Aggregate throughput against per-stream rate as concurrent streams increase, with TTFT p95 at each level. Aggregate rises while each stream slows; sizing a deployment needs both columns.

gpt-oss 20B · Q4_K_M — peaks at 862.1 tok/s aggregate
AGGREGATE TOK/S
0431862124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
0132265124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1264.6 tok/s264.6 tok/s31.7 ms32.2 ms200
2211.7 tok/s105.8 tok/s57.6 ms59.6 ms200
4322 tok/s80.5 tok/s77.9 ms96.3 ms200
8381 tok/s47.6 tok/s116.8 ms139.3 ms240
16862.1 tok/s53.9 tok/s132.6 ms149.9 ms480
Between 8 and 16 streams the aggregate rose 2.26× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 784.9 tok/s aggregate
AGGREGATE TOK/S
03927851248
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
01212421248
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1242 tok/s242 tok/s28.1 ms33.2 ms200
2399.6 tok/s199.8 tok/s31.5 ms64.1 ms200
4249.9 tok/s62.5 tok/s69 ms97.4 ms200
8784.9 tok/s98.1 tok/s82.5 ms99.8 ms240
Between 4 and 8 streams the aggregate rose 3.14× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,321.4 tok/s aggregate
AGGREGATE TOK/S
06611,321124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
0121241124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1241.1 tok/s241.1 tok/s12.9 ms16.4 ms200
2197.4 tok/s98.7 tok/s22.8 ms43.8 ms200
4350.8 tok/s87.7 tok/s42.9 ms99.4 ms200
8439.9 tok/s55 tok/s87.6 ms92.4 ms240
161,321.4 tok/s82.6 tok/s137 ms254 ms480
Between 8 and 16 streams the aggregate rose 3.00× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 1,135.9 tok/s aggregate
AGGREGATE TOK/S
05681,136124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
089.0178124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1177.9 tok/s177.9 tok/s16.3 ms17.7 ms200
2153.4 tok/s76.7 tok/s29.9 ms50.5 ms200
4267.8 tok/s67 tok/s43.8 ms90.9 ms200
8331.6 tok/s41.5 tok/s118.9 ms177.9 ms240
161,135.9 tok/s71 tok/s135.8 ms161 ms480
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 846.9 tok/s aggregate
AGGREGATE TOK/S
0423847124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
089.3179124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1178.5 tok/s178.5 tok/s35.8 ms36.5 ms200
2104.1 tok/s52.1 tok/s68.3 ms75.4 ms200
4192.7 tok/s48.2 tok/s82.1 ms106.5 ms200
8247.3 tok/s30.9 tok/s169.9 ms215.1 ms240
16846.9 tok/s52.9 tok/s141.6 ms149.4 ms480
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 1,189.4 tok/s aggregate
AGGREGATE TOK/S
05951,189124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
082.0164124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1164.1 tok/s164.1 tok/s14.7 ms15.8 ms200
2146.3 tok/s73.2 tok/s27.1 ms45.2 ms200
4268.4 tok/s67.1 tok/s52.6 ms98.2 ms200
8344.6 tok/s43.1 tok/s82.5 ms153.9 ms240
161,189.4 tok/s74.3 tok/s130.8 ms170.2 ms480
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 1,073.1 tok/s aggregate
AGGREGATE TOK/S
05371,073124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
078.1156124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1156.3 tok/s156.3 tok/s16.3 ms17.5 ms200
2136.5 tok/s68.2 tok/s29.5 ms48.8 ms200
4247.4 tok/s61.9 tok/s42.4 ms49.3 ms200
8319 tok/s39.9 tok/s98.5 ms132.1 ms240
161,073.1 tok/s67.1 tok/s130.9 ms140.5 ms480
Between 8 and 16 streams the aggregate rose 3.36× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 986 tok/s aggregate
AGGREGATE TOK/S
0493986124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
069.8140124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1139.7 tok/s139.7 tok/s16.5 ms18 ms200
2123.8 tok/s61.9 tok/s31 ms51.6 ms200
4228.5 tok/s57.1 tok/s58.4 ms89.5 ms200
8291.6 tok/s36.4 tok/s152.6 ms155.7 ms240
16986 tok/s61.6 tok/s189.3 ms264.3 ms480
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 707.3 tok/s aggregate
AGGREGATE TOK/S
0354707124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
068.4137124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1136.8 tok/s136.8 tok/s58.3 ms126.6 ms200
2116.2 tok/s58.1 tok/s127.3 ms192.6 ms200
4205.2 tok/s51.3 tok/s243.3 ms500.8 ms200
8247.9 tok/s31 tok/s467.8 ms736.5 ms240
16707.3 tok/s44.2 tok/s839.4 ms961.8 ms480
Between 8 and 16 streams the aggregate rose 2.85× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 939.5 tok/s aggregate
AGGREGATE TOK/S
0470939124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
062.1124124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1124.3 tok/s124.3 tok/s18.4 ms20.7 ms200
2111.6 tok/s55.8 tok/s33.5 ms53.1 ms200
4208.3 tok/s52.1 tok/s69.1 ms106.1 ms200
8269.1 tok/s33.6 tok/s120 ms144.1 ms240
16939.5 tok/s58.7 tok/s158 ms174.3 ms480
Between 8 and 16 streams the aggregate rose 3.49× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 895.7 tok/s aggregate
AGGREGATE TOK/S
0448896124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
050.6101124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1101.3 tok/s101.3 tok/s18.8 ms20.1 ms200
292.7 tok/s46.4 tok/s35 ms54.2 ms200
4174.6 tok/s43.7 tok/s45.1 ms71.4 ms200
8225.2 tok/s28.1 tok/s123 ms134.6 ms240
16895.7 tok/s56 tok/s133.1 ms139.8 ms480
Between 8 and 16 streams the aggregate rose 3.98× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 660.9 tok/s aggregate
AGGREGATE TOK/S
0330661124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
049.097.9124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
197.9 tok/s97.9 tok/s37.7 ms53.9 ms200
285.2 tok/s42.6 tok/s70.9 ms136 ms200
4156.3 tok/s39.1 tok/s100.6 ms353.8 ms200
8201.3 tok/s25.2 tok/s330.7 ms560.5 ms240
16660.9 tok/s41.3 tok/s304.4 ms523.7 ms480
Between 8 and 16 streams the aggregate rose 3.28× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 743.5 tok/s aggregate
AGGREGATE TOK/S
0372743124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
046.893.6124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
193.6 tok/s93.6 tok/s23.9 ms25.5 ms200
284.8 tok/s42.4 tok/s44.7 ms67.2 ms200
4159.3 tok/s39.8 tok/s65.3 ms75.2 ms200
8207.3 tok/s25.9 tok/s134.2 ms181.9 ms240
16743.5 tok/s46.5 tok/s211 ms236.8 ms480
Between 8 and 16 streams the aggregate rose 3.59× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 639.3 tok/s aggregate
AGGREGATE TOK/S
0320639124816
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
044.088.0124816
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
188 tok/s88 tok/s39.7 ms54.2 ms200
278.5 tok/s39.2 tok/s75.2 ms145.3 ms200
4143.7 tok/s35.9 tok/s167.2 ms367.1 ms200
8185.4 tok/s23.2 tok/s309.9 ms506.7 ms240
16639.3 tok/s40 tok/s310 ms585.2 ms480
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 356.3 tok/s aggregate
AGGREGATE TOK/S
01783561248
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
032.164.21248
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
164.2 tok/s64.2 tok/s46.2 ms65.4 ms200
2119.9 tok/s59.9 tok/s93.1 ms121.5 ms200
4110.5 tok/s27.6 tok/s173.8 ms303 ms200
8356.3 tok/s44.5 tok/s295 ms429.8 ms240
Between 4 and 8 streams the aggregate rose 3.22× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 292.1 tok/s aggregate
AGGREGATE TOK/S
01462921248
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
030.961.91248
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
161.9 tok/s61.9 tok/s31.1 ms39.2 ms200
2117.1 tok/s58.5 tok/s51.4 ms96.2 ms200
4111.2 tok/s27.8 tok/s124.7 ms127 ms200
8292.1 tok/s36.5 tok/s205.2 ms212.9 ms240
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Gemma 3 27B Instruct · Q4_K_M — peaks at 221.7 tok/s aggregate
AGGREGATE TOK/S
01112221248
CONCURRENT STREAMS
TOK/S EACH STREAM SEES
024.348.61248
CONCURRENT STREAMS
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
148.6 tok/s48.6 tok/s65.6 ms90.5 ms200
290.1 tok/s45.1 tok/s94.8 ms180 ms200
484.3 tok/s21.1 tok/s327.8 ms463.9 ms200
8221.7 tok/s27.7 tok/s375.8 ms449 ms240
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

ONE FLAG

Flash attention, measured at depth

Flash attention on and off, at depth. LM Studio and Ollama enable it by default. The effect depends on the model and on cache size, so it is measured with a filled KV cache rather than an empty one.

gpt-oss 20B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off250.74,999.911.8 GB3,968
on272.18,532.111.4 GB3,968
Qwen3 30B A3B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off197.33,146.218.4 GB3,968
on218.95,947.818.3 GB3,968
Qwen3 4B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off204.94,501.93.8 GB3,968
on2209,815.43.7 GB3,968
GLM 4.7 Flash · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off95.22,856.517.9 GB3,968
on167.33,970.817.9 GB3,968
Llama 3.1 8B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off1484,432.85.7 GB3,968
on155.18,314.75.5 GB3,968
Qwen3 8B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off140.24,033.95.8 GB3,968
on147.37,762.85.7 GB3,968
Qwen3.5 9B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off138.86,382.95.9 GB3,968
on140.77,066.95.9 GB3,968
Gemma 3 12B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off92.94,623.68.5 GB3,968
on95.36,008.28.5 GB3,968
Qwen3 14B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off85.32,726.99.6 GB3,968
on88.84,645.89.4 GB3,968
Mistral Small 3.2 24B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off58.42,424.514.5 GB3,968
on59.73,535.214.4 GB3,968
Gemma 3 27B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off46.42,330.717.4 GB3,968
on47.22,973.317.3 GB3,968
WATTS AND MONEY

Power and cost per million tokens

Power sampled at the card during generation, not the TDP from the spec sheet. Cost per million tokens uses the hourly rate this machine was billed at; for owned hardware, apply your own electricity price to the same energy figures.

gpt-oss 20B · Q4_K_M0.664
Qwen3 30B A3B Instruct 2507 · Q4_K_M0.740
Qwen3 4B Instruct 2507 · Q4_K_M0.742
Qwen3 8B (Q3_K_M)1.03
GLM 4.7 Flash · Q4_K_M1.03
Llama 3.1 8B Instruct · Q4_K_M1.12
Qwen3 8B · Q4_K_M1.17
Qwen3 8B (Q5_K_M)1.32
Qwen3.5 9B · Q4_K_M1.33
Qwen3 8B (Q6_K)1.49
Qwen3 8B (Q8_0)1.85
Gemma 3 12B Instruct · Q4_K_M1.88
Qwen3 14B · Q4_K_M2.00
Gemma 3 12B Instruct (Q5_K_M)2.11
Gemma 3 12B Instruct (Q8_0)2.92
Mistral Small 3.2 24B Instruct · Q4_K_M3.07
Gemma 3 27B Instruct · Q4_K_M3.88
$ PER MILLION TOKENS, ONE STREAM — CHEAPER IS SHORTER
ModelQuantAvg WPeak WPeak °Ctok/WkWh/Mtok$/Mtok single$/Mtok batched
gpt-oss 20BQ4_K_M210338471.380.202$0.66$0.22
Qwen3 30B A3B Instruct 2507Q4_K_M181307451.430.195$0.74$0.24
Qwen3 4B Instruct 2507Q4_K_M268332490.970.288$0.74$0.15
Qwen3 8B (Q3_K_M)Q3_K_M338414550.550.504$1.03$0.17
GLM 4.7 FlashQ4_K_M1873024410.278$1.03$0.23
Llama 3.1 8B InstructQ4_K_M299361490.570.486$1.12$0.16
Qwen3 8BQ4_K_M298356510.550.506$1.17$0.18
Qwen3 8B (Q5_K_M)Q5_K_M298359520.490.572$1.32$0.19
Qwen3.5 9BQ4_K_M291357520.50.562$1.33$0.27
Qwen3 8B (Q6_K)Q6_K326390570.390.707$1.49$0.2
Qwen3 8B (Q8_0)Q8_0267314490.390.715$1.85$0.21
Gemma 3 12B InstructQ4_K_M302366510.340.823$1.88$0.29
Qwen3 14BQ4_K_M316379530.30.915$2$0.26
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M305362530.30.932$2.11$0.3
Gemma 3 12B Instruct (Q8_0)Q8_0266329530.251.129$2.92$0.54
Mistral Small 3.2 24B InstructQ4_K_M329413590.191.461$3.07$0.66
Gemma 3 27B InstructQ4_K_M323409590.151.815$3.88$0.86
THE RUN, TAKEN TOGETHER

Relationships across the whole run

Relationships fitted across the whole run rather than measurements of one model: the card's decode constants, the runtime's fixed VRAM overhead, and how TTFT scales with concurrency. Each states its n and its correlation coefficient, and a fit with a weak correlation is quoted without a line drawn through it.

DECODE CONSTANTS: FIXED COST AND BANDWIDTH

At batch one, decode reads every weight once per token. Time per token is therefore a fixed cost plus weight bytes divided by bandwidth, which is a straight line when milliseconds per token is plotted against gigabytes. Fitted over 5 quantisations of Qwen3 8B. r = 1.000 across 5 models.

MILLISECONDS PER TOKEN
0.04.89.604.068.11(Q8_0)(Q6_K)(Q5_K_M)Q4_K_M(Q3_K_M)
WEIGHTS, GB
IMPLIED BANDWIDTH
989 GB/s
inverse of the slope
FIXED PER TOKEN
1.43 ms
intercept; not explained by size
SO A 14 GB MODEL
64.1 tok/s
predicted, not measured
An implied bandwidth above the card’s rated figure does not mean the card is faster than rated. It means the file size on disk is not exactly the number of bytes the runtime moved per token. Treat the two constants as the pair that reproduces these measurements. The 14 GB figure is what they predict; no 14 GB model was measured in this run.
FIXED VRAM OVERHEAD BEFORE THE WEIGHTS

Peak VRAM against weight size fits a line of slope near one. The intercept is everything that is not weights: KV cache, compute buffers and the allocator’s own reservation. A fit estimate that assumes a zero intercept will report that a model fits when it does not. r = 0.991 across 17 models.

PEAK VRAM, GB
09.0118.008.6417.3Qwen3 30B A3B…Gemma 3 27B In…Mistral Small…Gemma 3 12B In…gpt-oss 20BGemma 3 12B In…Qwen3 14BGemma 3 12B In…Qwen3 8B (Q6_K)Qwen3 8B (Q5_K…Qwen3.5 9BQwen3 8BQwen3 4B Instr…3142
WEIGHTS, GB
1 GLM 4.7 Flash2 Qwen3 8B (Q8_0)3 Llama 3.1 8B I…4 Qwen3 8B (Q3_K…
FIXED OVERHEAD
0.8 GB
intercept, weights excluded
PER GB OF WEIGHTS
1.03 GB
slope; one would be exact
TTFT P95 AGAINST CONCURRENT STREAMS

Time to first token at the 95th percentile against concurrent streams, on Llama 3.1 8B Instruct · Q4_K_M. Aggregate throughput does not show this: it rises while each individual request waits longer. r = 0.89 across 5 models: a relationship, with visible scatter.

TTFT P95, MS
085.117008.0016.0168421
CONCURRENT STREAMS
EACH ADDED STREAM
+10 ms
onto the 95th percentile
AT ONE STREAM
36 ms
intercept, one stream
THE MACHINE

What this ran on, in full

The full hardware and software configuration these numbers were taken on. Fields that require privileges the harness did not have — the DMI table is not readable inside a container — are marked unavailable rather than omitted, since a missing field and an unreadable one are different.

PropertyValue
CPUAMD EPYC 7452 32-Core Processor
Cores32 physical / 64 logical
Cores this process could use10
Threads given to llama-bench10
NUMA nodes1
RAM504 GB
Kernel6.8.0-124-generic
OSUbuntu 24.04.1 LTS
SystemTo Be Filled By O.E.M. ROMED8-2T/BCM
BIOSP4.10 06/05/2025
Containerisedyes
Disksnvme0n1 Samsung SSD 980 PRO 250GB 250 GB, nvme1n1 SAMSUNG MZQL27T6HBLA-00A07 7682 GB
GPUVRAMPCIe linkPower limitMax SM / mem clock
0: NVIDIA GeForce RTX 409024.0 GBgen 1 ×16 (of gen 4 ×16)450 W3,135 / 10,501 MHz
▸ WANT ONE FOR YOUR MACHINE?

The NVIDIA GeForce RTX 4090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.

You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.

ASK ABOUT A REPORT →
fitmyllm@gmail.com
HOW THIS WAS PRODUCED

A pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-4090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →