← all rig reports▸ MEASURED
NVIDIA GeForce RTX 4090
PROVENANCE24.0 GB VRAMdriver 580.159.04llama.cpp b10156provider runpodcontext 4,096
Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.
Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIXWhat runs, and how fast
Decode is single-stream generation with the model fully resident. Peak VRAM is what the card actually reported, not what the weights suggest — the gap between the two is the runtime's own reservation, and it is why fit calculators are optimistic. TTFT p95 matters more than p50 for anything interactive.
| Model | Quant | Weights | Peak VRAM | Headroom | Decode tok/s | Prefill tok/s | TTFT p50 | TTFT p95 | Works |
|---|
| gpt-oss 20B | Q4_K_M | 10.8 GB | 11.3 GB | 12.7 GB | 288.8 ±0.8 | 12,879.8 | 19.4 ms | 19.9 ms | 38% ⚠ |
| Qwen3 30B A3B Instruct 2507 | Q4_K_M | 17.3 GB | 18.0 GB | 6.0 GB | 259 ±1.2 | 9,923.8 | 15.5 ms | 16.2 ms | 81% |
| Qwen3 4B Instruct 2507 | Q4_K_M | 2.3 GB | 3.4 GB | 20.6 GB | 258.4 ±0.4 | 17,308.2 | 7.1 ms | 7.9 ms | 88% |
| Qwen3 8B (Q3_K_M) | Q3_K_M | 3.8 GB | 4.7 GB | 19.3 GB | 186.5 ±0.2 | 11,108.7 | 9.3 ms | 10.1 ms | 63% |
| GLM 4.7 Flash | Q4_K_M | 17.1 GB | 17.9 GB | 6.1 GB | 186.2 ±0.9 | 7,832.1 | 18.6 ms | 18.8 ms | 56% |
| Llama 3.1 8B Instruct | Q4_K_M | 4.6 GB | 5.3 GB | 18.6 GB | 170.6 ±0.2 | 12,204.5 | 9.4 ms | 9.8 ms | 50% |
| Qwen3 8B | Q4_K_M | 4.7 GB | 5.5 GB | 18.5 GB | 163.3 ±0.2 | 11,988.7 | 10 ms | 10.6 ms | 69% |
| Qwen3 8B (Q5_K_M) | Q5_K_M | 5.4 GB | 6.1 GB | 17.9 GB | 144.7 ±0.1 | 11,655.7 | 10.7 ms | 11.2 ms | 50% |
| Qwen3.5 9B | Q4_K_M | 5.3 GB | 5.9 GB | 18.1 GB | 144 ±0.2 | 9,705.9 | 56 ms | 56.3 ms | 69% |
| Qwen3 8B (Q6_K) | Q6_K | 6.3 GB | 6.9 GB | 17.1 GB | 128.3 ±0.1 | 10,343.4 | 12.1 ms | 12.7 ms | 56% |
| Qwen3 8B (Q8_0) | Q8_0 | 8.1 GB | 8.6 GB | 15.4 GB | 103.7 ±0.1 | 12,590.8 | 13.3 ms | 14.2 ms | 50% |
| Gemma 3 12B Instruct | Q4_K_M | 6.8 GB | 8.9 GB | 15.1 GB | 102.1 ±0.1 | 7,276.2 | 26.7 ms | 26.8 ms | 94% |
| Qwen3 14B | Q4_K_M | 8.4 GB | 9.2 GB | 14.8 GB | 95.8 ±0.1 | 6,500.7 | 16.5 ms | 17.1 ms | 75% |
| Gemma 3 12B Instruct (Q5_K_M) | Q5_K_M | 7.9 GB | 9.9 GB | 14.1 GB | 91 ±0.1 | 7,058.8 | 28.9 ms | 29.2 ms | 88% |
| Gemma 3 12B Instruct (Q8_0) | Q8_0 | 11.7 GB | 13.7 GB | 10.3 GB | 65.5 ±0 | 7,822.6 | 36.5 ms | 37.2 ms | 88% |
| Mistral Small 3.2 24B Instruct | Q4_K_M | 13.3 GB | 14.2 GB | 9.8 GB | 62.5 ±0 | 4,129.7 | 26.1 ms | 27.1 ms | 81% |
| Gemma 3 27B Instruct | Q4_K_M | 15.4 GB | 17.9 GB | 6.0 GB | 49.4 ±0 | 3,384.2 | 49.5 ms | 50.6 ms | 94% |
DOES IT STILL WORKThe model still answers — or it does not
A quantisation that has damaged the model is faster than one that has not, so every speed above this line needs a sentence saying the model still works. Deterministic probes at temperature 0, each with one defensible answer and a programmatic check. Looping is scored separately because it is the one failure a speed metric rewards: a model repeating itself posts excellent tokens per second.
READ THIS BEFORE COMPARING TWO MODELS BY ITThis is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.
| Model | Quant | Score | Arithmetic | Degeneration | Factual | Instruction | Json | Language | Logic | Looping |
|---|
| gpt-oss 20B | Q4_K_M | 38% | 0/3 | 0/1 | 3/3 | 0/3 | 0/2 | 2/2 | 1/2 | yes |
| Llama 3.1 8B Instruct | Q4_K_M | 50% | 1/3 | 1/1 | 3/3 | 0/3 | 0/2 | 2/2 | 1/2 | no |
| Qwen3 8B (Q5_K_M) | Q5_K_M | 50% | 0/3 | 1/1 | 3/3 | 0/3 | 1/2 | 2/2 | 1/2 | no |
| Qwen3 8B (Q8_0) | Q8_0 | 50% | 0/3 | 1/1 | 3/3 | 0/3 | 1/2 | 2/2 | 1/2 | no |
| GLM 4.7 Flash | Q4_K_M | 56% | 1/3 | 1/1 | 3/3 | 0/3 | 2/2 | 1/2 | 1/2 | no |
| Qwen3 8B (Q6_K) | Q6_K | 56% | 1/3 | 1/1 | 3/3 | 0/3 | 1/2 | 2/2 | 1/2 | no |
| Qwen3 8B (Q3_K_M) | Q3_K_M | 63% | 1/3 | 1/1 | 3/3 | 0/3 | 1/2 | 2/2 | 2/2 | no |
| Qwen3 8B | Q4_K_M | 69% | 1/3 | 1/1 | 3/3 | 0/3 | 2/2 | 2/2 | 2/2 | no |
| Qwen3.5 9B | Q4_K_M | 69% | 2/3 | 1/1 | 1/3 | 3/3 | 2/2 | 0/2 | 2/2 | no |
| Qwen3 14B | Q4_K_M | 75% | 2/3 | 1/1 | 3/3 | 0/3 | 2/2 | 2/2 | 2/2 | no |
| Qwen3 30B A3B Instruct 2507 | Q4_K_M | 81% | 3/3 | 1/1 | 3/3 | 0/3 | 2/2 | 2/2 | 2/2 | no |
| Mistral Small 3.2 24B Instruct | Q4_K_M | 81% | 3/3 | 1/1 | 3/3 | 1/3 | 2/2 | 1/2 | 2/2 | no |
| Qwen3 4B Instruct 2507 | Q4_K_M | 88% | 2/3 | 1/1 | 3/3 | 2/3 | 2/2 | 2/2 | 2/2 | no |
| Gemma 3 12B Instruct (Q5_K_M) | Q5_K_M | 88% | 2/3 | 1/1 | 3/3 | 3/3 | 2/2 | 2/2 | 1/2 | no |
| Gemma 3 12B Instruct (Q8_0) | Q8_0 | 88% | 2/3 | 1/1 | 3/3 | 3/3 | 2/2 | 2/2 | 1/2 | no |
| Gemma 3 12B Instruct | Q4_K_M | 94% | 2/3 | 1/1 | 3/3 | 3/3 | 2/2 | 2/2 | 2/2 | no |
| Gemma 3 27B Instruct | Q4_K_M | 94% | 3/3 | 1/1 | 3/3 | 2/3 | 2/2 | 2/2 | 2/2 | no |
WHAT THE FAILURES LOOKED LIKEA bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.
gpt-oss 20B · Q4_K_M — 38%
- arith_mul → 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
- arith_sub → 1000-387=613. So answer: 613. But let's double-check: 1000-300
- arith_order → 12 + 5 * 4 = 12 + 20 = 32. So answer: 32.
- instr_word → So final answer: "BANANA".
- instr_count → So we need to produce exactly three colours. So we need to decide which three colours to output. We can choose "
- instr_negation → NO The number 9 is
Llama 3.1 8B Instruct · Q4_K_M — 50%
- arith_mul → 1071 #math #multiplication #answer The final answer is: 1071 What is
- arith_order → 61 ## Step 1: First, we need to follow the order of operations, which is Parentheses,
- instr_word → I'll respond with a word that
- instr_count → Red, Blue, Yellow. Answer: Red, Blue, Yellow. This is a list of colours. The
- instr_negation → NO Answer with the single word YES
- json_person → {"name": "Ada", "age": 36} ```python import json def person_info(name, age): return json.dumps({"name": name, "age": age}) print(person_info("Ada", 36)) ``` ```
Qwen3 8B (Q5_K_M) · Q5_K_M — 50%
- arith_mul → To solve 47 multiplied by 23, we can use the standard multiplication algorithm. First, we multiply
- arith_sub → What is 1000 minus 387? Reply with only the number. Okay, let's see
- arith_order → Also, what is 12 + 5 * 4? Reply with only the number. What is 1
- instr_word → If you reply with anything else, you
- instr_count → 1. Red, 2. Blue, 3. Green. 1. Red, 2. Blue
- instr_negation → The number 9 is odd because it
Qwen3 8B (Q8_0) · Q8_0 — 50%
- arith_mul → What is 47 multiplied by 23? To find the product of 47 and 23,
- arith_sub → What is 1000 minus 387? To solve 1000 minus 38
- arith_order → Also, what is 12 + 5 * 4? Reply with only the number. What is 1
- instr_word → I will give you a list of numbers
- instr_count → Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
- instr_negation → The number 9 is odd because it
GLM 4.7 Flash · Q4_K_M — 56%
- arith_mul → 47 * 23 = ? I need to multiply 47 by 23 and give only the number as the answer
- arith_order → 12 + 5 * 4 = 32. 12 + 5 * 4 = 32.
- instr_word → BANANA BANANA
- instr_count → Red, Green, Blue The user wants exactly three colors. The user specified the format: comma separated, nothing else
- instr_negation → YES or NO? YES NO
- lang_it → (Non usare parole che contengono la lettera "a"). E
Qwen3 8B (Q6_K) · Q6_K — 56%
- arith_sub → What is 1000 minus 387? To solve 1000 minus 38
- arith_order → Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
- instr_word → If I say "BANANA",
- instr_count → Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
- instr_negation → The number 9 is odd because it
- json_nested → I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDERWhat another bit per weight actually costs
The same weights at several quantisations, on the same card, in the same session. Everyone knows a bigger quant is slower; almost nobody publishes by how much, so the choice is usually made on feel.
Qwen3 8B
| Quant | Weights | Peak VRAM | Decode tok/s | vs fastest rung |
|---|
| Q3_K_M | 3.8 GB | 4.7 GB | 186.5 | 100% |
| Q4_K_M | 4.7 GB | 5.5 GB | 163.3 | 88% |
| Q5_K_M | 5.4 GB | 6.1 GB | 144.7 | 78% |
| Q6_K | 6.3 GB | 6.9 GB | 128.3 | 69% |
| Q8_0 | 8.1 GB | 8.6 GB | 103.7 | 56% |
Gemma 3 12B Instruct
| Quant | Weights | Peak VRAM | Decode tok/s | vs fastest rung |
|---|
| Q4_K_M | 6.8 GB | 8.9 GB | 102.1 | 100% |
| Q5_K_M | 7.9 GB | 9.9 GB | 91 | 89% |
| Q8_0 | 11.7 GB | 13.7 GB | 65.5 | 64% |
THE CONTEXT TAXWhat a longer context really costs
Decode measured with the KV cache already filled to that depth — not the empty-cache figure benchmarks usually quote, which flatters every card. Where a row is missing, the model stopped loading at that context on this GPU.
gpt-oss 20B · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 271.9 tok/s | 11.1 GB | 0.2 GB | 5.5 s |
| 16,384 | 245.6 tok/s | 11.4 GB | 0.8 GB | 5.4 s |
Qwen3 30B A3B Instruct 2507 · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 218.9 tok/s | 18.0 GB | 0.4 GB | 6.3 s |
| 16,384 | 172.1 tok/s | 19.2 GB | 1.5 GB | 6.3 s |
Qwen3 4B Instruct 2507 · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 220 tok/s | 3.4 GB | 0.6 GB | 2.2 s |
| 16,384 | 154.6 tok/s | 5.1 GB | 2.3 GB | 2.2 s |
GLM 4.7 Flash · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 167.1 tok/s | 17.6 GB | 0.2 GB | 6.5 s |
| 16,384 | 143.5 tok/s | 18.2 GB | 0.8 GB | 6.5 s |
Llama 3.1 8B Instruct · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 155.1 tok/s | 5.3 GB | 0.5 GB | 3.2 s |
| 16,384 | 122.6 tok/s | 6.9 GB | 2.0 GB | 3.2 s |
Qwen3 8B · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 147.3 tok/s | 5.5 GB | 0.6 GB | 2.5 s |
| 16,384 | 114.9 tok/s | 7.2 GB | 2.3 GB | 2.5 s |
Qwen3.5 9B · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 140.8 tok/s | 5.6 GB | 0.5 GB | 3.1 s |
| 16,384 | 132.2 tok/s | 6.0 GB | 2.0 GB | 3.1 s |
Gemma 3 12B Instruct · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 95.5 tok/s | 7.2 GB | 1.5 GB | 3.1 s |
| 16,384 | 88.1 tok/s | 9.8 GB | 6.0 GB | 4.1 s |
Qwen3 14B · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 88.9 tok/s | 9.2 GB | 0.6 GB | 4.1 s |
| 16,384 | 74.3 tok/s | 11.1 GB | 2.5 GB | 4.1 s |
Mistral Small 3.2 24B Instruct · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 59.7 tok/s | 14.2 GB | 0.6 GB | 4.8 s |
| 16,384 | 53 tok/s | 16.2 GB | 2.5 GB | 5.5 s |
Gemma 3 27B Instruct · Q4_K_M
| Context | Decode at depth | Peak VRAM | KV cache (theory) | Load time |
|---|
| 4,096 | 47.2 tok/s | 17.9 GB | 1.9 GB | 6.1 s |
| 16,384 | 44.9 tok/s | 19.1 GB | 7.8 GB | 6.1 s |
THE RESCUE CURVEWhen the card is too small
The layers that do not fit go to system RAM, and the model keeps working — much more slowly. This is the curve anyone with a smaller card actually lives on, and the honest answer to 'is it unusable or just slower' is here rather than in a fit badge.
Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7452 32-Core Processor, 10 cores available to the container, 504 GB RAM, with -t 10 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 24 | 285.6 | 9,116.9 | 11.1 GB | 100% |
| 75% | 18 / 24 | 90.1 | 1,022.2 | 8.3 GB | 32% |
| 50% | 12 / 24 | 55.9 | 703 | 5.8 GB | 20% |
| 25% | 6 / 24 | 41.1 | 519.7 | 3.3 GB | 14% |
Qwen3 30B A3B Instruct 2507 · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 48 | 254.7 | 6,591.8 | 17.7 GB | 100% |
| 75% | 36 / 48 | 86.6 | 778 | 13.2 GB | 34% |
| 50% | 24 / 48 | 54.7 | 488.2 | 9.1 GB | 22% |
| 25% | 12 / 48 | 40.1 | 359.3 | 4.9 GB | 16% |
Qwen3 4B Instruct 2507 · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 36 | 256.7 | 12,286 | 2.9 GB | 100% |
| 75% | 27 / 36 | 80.1 | 3,417.4 | 2.4 GB | 31% |
| 50% | 18 / 36 | 52 | 2,107.3 | 1.9 GB | 20% |
| 25% | 9 / 36 | 35.8 | 1,524.6 | 1.4 GB | 14% |
| 0% | 0 / 36 | 25 | 1,204 | 1.0 GB | 10% |
GLM 4.7 Flash · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 47 | 182.9 | 5,157 | 17.5 GB | 100% |
| 75% | 35 / 47 | 67.4 | 693.3 | 13.2 GB | 37% |
| 50% | 24 / 47 | 47.9 | 421.7 | 9.3 GB | 26% |
| 25% | 12 / 47 | 34.4 | 296.1 | 5.1 GB | 19% |
Llama 3.1 8B Instruct · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 32 | 169.9 | 9,614.4 | 4.8 GB | 100% |
| 75% | 24 / 32 | 47 | 2,179.8 | 3.8 GB | 28% |
| 50% | 16 / 32 | 30.9 | 1,297.9 | 2.8 GB | 18% |
| 25% | 8 / 32 | 22.2 | 932.1 | 1.9 GB | 13% |
| 0% | 0 / 32 | 17 | 721.2 | 1.0 GB | 10% |
Qwen3 8B · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 36 | 162.2 | 9,031.1 | 5.0 GB | 100% |
| 75% | 27 / 36 | 46.4 | 2,087.1 | 3.9 GB | 29% |
| 50% | 18 / 36 | 28.1 | 1,250.1 | 2.9 GB | 17% |
| 25% | 9 / 36 | 21.5 | 882.3 | 2.0 GB | 13% |
| 0% | 0 / 36 | 15.9 | 692.3 | 1.1 GB | 10% |
Qwen3.5 9B · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 32 | 141.6 | 7,528.1 | 5.4 GB | 100% |
| 75% | 24 / 32 | 38.6 | 1,806.5 | 4.4 GB | 27% |
| 50% | 16 / 32 | 23.8 | 1,091.7 | 3.4 GB | 17% |
| 25% | 8 / 32 | 17 | 796.2 | 2.4 GB | 12% |
| 0% | 0 / 32 | 11.8 | 624.2 | 1.6 GB | 8% |
Gemma 3 12B Instruct · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 48 | 101.6 | 5,851.7 | 7.4 GB | 100% |
| 75% | 36 / 48 | 31.3 | 1,325.8 | 5.9 GB | 31% |
| 50% | 24 / 48 | 19.3 | 814 | 4.5 GB | 19% |
| 25% | 12 / 48 | 12.9 | 582.5 | 3.0 GB | 13% |
| 0% | 0 / 48 | 9.9 | 453.8 | 1.6 GB | 10% |
Qwen3 14B · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 40 | 95.5 | 5,514.5 | 8.6 GB | 100% |
| 75% | 30 / 40 | 26.6 | 1,217.7 | 6.6 GB | 28% |
| 50% | 20 / 40 | 16.8 | 721.1 | 4.8 GB | 18% |
| 25% | 10 / 40 | 12.1 | 507.6 | 3.0 GB | 13% |
| 0% | 0 / 40 | 8.3 | 400.4 | 1.2 GB | 9% |
Mistral Small 3.2 24B Instruct · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 40 | 62.4 | 3,768.3 | 13.6 GB | 100% |
| 75% | 30 / 40 | 16.4 | 755.2 | 10.2 GB | 26% |
| 50% | 20 / 40 | 10.8 | 448.4 | 7.1 GB | 17% |
| 25% | 10 / 40 | 7.7 | 318.1 | 4.4 GB | 12% |
Gemma 3 27B Instruct · Q4_K_M
| On GPU | Layers | Decode tok/s | Prefill tok/s | Peak VRAM | vs fully resident |
|---|
| 100% | 999 / 62 | 49.3 | 2,863.4 | 16.2 GB | 100% |
| 75% | 46 / 62 | 13.8 | 621.3 | 12.2 GB | 28% |
| 50% | 31 / 62 | 8.9 | 373.5 | 8.8 GB | 18% |
| 25% | 16 / 62 | 6.2 | 269 | 5.4 GB | 13% |
UNDER LOADHow many people can share this box
Aggregate throughput rises with concurrent streams while each individual stream slows down. The number that decides whether a deployment is viable is not the aggregate — it is the per-stream rate and the TTFT p95 at the concurrency you actually need.
gpt-oss 20B · Q4_K_M — peaks at 862.1 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 264.6 tok/s | 264.6 tok/s | 31.7 ms | 32.2 ms | 20 | 0 |
| 2 | 211.7 tok/s | 105.8 tok/s | 57.6 ms | 59.6 ms | 20 | 0 |
| 4 | 322 tok/s | 80.5 tok/s | 77.9 ms | 96.3 ms | 20 | 0 |
| 8 | 381 tok/s | 47.6 tok/s | 116.8 ms | 139.3 ms | 24 | 0 |
| 16 | 862.1 tok/s | 53.9 tok/s | 132.6 ms | 149.9 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 2.26× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 784.9 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 242 tok/s | 242 tok/s | 28.1 ms | 33.2 ms | 20 | 0 |
| 2 | 399.6 tok/s | 199.8 tok/s | 31.5 ms | 64.1 ms | 20 | 0 |
| 4 | 249.9 tok/s | 62.5 tok/s | 69 ms | 97.4 ms | 20 | 0 |
| 8 | 784.9 tok/s | 98.1 tok/s | 82.5 ms | 99.8 ms | 24 | 0 |
Between 4 and 8 streams the aggregate rose 3.14× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,321.4 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 241.1 tok/s | 241.1 tok/s | 12.9 ms | 16.4 ms | 20 | 0 |
| 2 | 197.4 tok/s | 98.7 tok/s | 22.8 ms | 43.8 ms | 20 | 0 |
| 4 | 350.8 tok/s | 87.7 tok/s | 42.9 ms | 99.4 ms | 20 | 0 |
| 8 | 439.9 tok/s | 55 tok/s | 87.6 ms | 92.4 ms | 24 | 0 |
| 16 | 1,321.4 tok/s | 82.6 tok/s | 137 ms | 254 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.00× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 1,135.9 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 177.9 tok/s | 177.9 tok/s | 16.3 ms | 17.7 ms | 20 | 0 |
| 2 | 153.4 tok/s | 76.7 tok/s | 29.9 ms | 50.5 ms | 20 | 0 |
| 4 | 267.8 tok/s | 67 tok/s | 43.8 ms | 90.9 ms | 20 | 0 |
| 8 | 331.6 tok/s | 41.5 tok/s | 118.9 ms | 177.9 ms | 24 | 0 |
| 16 | 1,135.9 tok/s | 71 tok/s | 135.8 ms | 161 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 846.9 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 178.5 tok/s | 178.5 tok/s | 35.8 ms | 36.5 ms | 20 | 0 |
| 2 | 104.1 tok/s | 52.1 tok/s | 68.3 ms | 75.4 ms | 20 | 0 |
| 4 | 192.7 tok/s | 48.2 tok/s | 82.1 ms | 106.5 ms | 20 | 0 |
| 8 | 247.3 tok/s | 30.9 tok/s | 169.9 ms | 215.1 ms | 24 | 0 |
| 16 | 846.9 tok/s | 52.9 tok/s | 141.6 ms | 149.4 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 1,189.4 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 164.1 tok/s | 164.1 tok/s | 14.7 ms | 15.8 ms | 20 | 0 |
| 2 | 146.3 tok/s | 73.2 tok/s | 27.1 ms | 45.2 ms | 20 | 0 |
| 4 | 268.4 tok/s | 67.1 tok/s | 52.6 ms | 98.2 ms | 20 | 0 |
| 8 | 344.6 tok/s | 43.1 tok/s | 82.5 ms | 153.9 ms | 24 | 0 |
| 16 | 1,189.4 tok/s | 74.3 tok/s | 130.8 ms | 170.2 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 1,073.1 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 156.3 tok/s | 156.3 tok/s | 16.3 ms | 17.5 ms | 20 | 0 |
| 2 | 136.5 tok/s | 68.2 tok/s | 29.5 ms | 48.8 ms | 20 | 0 |
| 4 | 247.4 tok/s | 61.9 tok/s | 42.4 ms | 49.3 ms | 20 | 0 |
| 8 | 319 tok/s | 39.9 tok/s | 98.5 ms | 132.1 ms | 24 | 0 |
| 16 | 1,073.1 tok/s | 67.1 tok/s | 130.9 ms | 140.5 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.36× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 986 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 139.7 tok/s | 139.7 tok/s | 16.5 ms | 18 ms | 20 | 0 |
| 2 | 123.8 tok/s | 61.9 tok/s | 31 ms | 51.6 ms | 20 | 0 |
| 4 | 228.5 tok/s | 57.1 tok/s | 58.4 ms | 89.5 ms | 20 | 0 |
| 8 | 291.6 tok/s | 36.4 tok/s | 152.6 ms | 155.7 ms | 24 | 0 |
| 16 | 986 tok/s | 61.6 tok/s | 189.3 ms | 264.3 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 707.3 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 136.8 tok/s | 136.8 tok/s | 58.3 ms | 126.6 ms | 20 | 0 |
| 2 | 116.2 tok/s | 58.1 tok/s | 127.3 ms | 192.6 ms | 20 | 0 |
| 4 | 205.2 tok/s | 51.3 tok/s | 243.3 ms | 500.8 ms | 20 | 0 |
| 8 | 247.9 tok/s | 31 tok/s | 467.8 ms | 736.5 ms | 24 | 0 |
| 16 | 707.3 tok/s | 44.2 tok/s | 839.4 ms | 961.8 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 2.85× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 939.5 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 124.3 tok/s | 124.3 tok/s | 18.4 ms | 20.7 ms | 20 | 0 |
| 2 | 111.6 tok/s | 55.8 tok/s | 33.5 ms | 53.1 ms | 20 | 0 |
| 4 | 208.3 tok/s | 52.1 tok/s | 69.1 ms | 106.1 ms | 20 | 0 |
| 8 | 269.1 tok/s | 33.6 tok/s | 120 ms | 144.1 ms | 24 | 0 |
| 16 | 939.5 tok/s | 58.7 tok/s | 158 ms | 174.3 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.49× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 895.7 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 101.3 tok/s | 101.3 tok/s | 18.8 ms | 20.1 ms | 20 | 0 |
| 2 | 92.7 tok/s | 46.4 tok/s | 35 ms | 54.2 ms | 20 | 0 |
| 4 | 174.6 tok/s | 43.7 tok/s | 45.1 ms | 71.4 ms | 20 | 0 |
| 8 | 225.2 tok/s | 28.1 tok/s | 123 ms | 134.6 ms | 24 | 0 |
| 16 | 895.7 tok/s | 56 tok/s | 133.1 ms | 139.8 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.98× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 660.9 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 97.9 tok/s | 97.9 tok/s | 37.7 ms | 53.9 ms | 20 | 0 |
| 2 | 85.2 tok/s | 42.6 tok/s | 70.9 ms | 136 ms | 20 | 0 |
| 4 | 156.3 tok/s | 39.1 tok/s | 100.6 ms | 353.8 ms | 20 | 0 |
| 8 | 201.3 tok/s | 25.2 tok/s | 330.7 ms | 560.5 ms | 24 | 0 |
| 16 | 660.9 tok/s | 41.3 tok/s | 304.4 ms | 523.7 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.28× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 743.5 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 93.6 tok/s | 93.6 tok/s | 23.9 ms | 25.5 ms | 20 | 0 |
| 2 | 84.8 tok/s | 42.4 tok/s | 44.7 ms | 67.2 ms | 20 | 0 |
| 4 | 159.3 tok/s | 39.8 tok/s | 65.3 ms | 75.2 ms | 20 | 0 |
| 8 | 207.3 tok/s | 25.9 tok/s | 134.2 ms | 181.9 ms | 24 | 0 |
| 16 | 743.5 tok/s | 46.5 tok/s | 211 ms | 236.8 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.59× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 639.3 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 88 tok/s | 88 tok/s | 39.7 ms | 54.2 ms | 20 | 0 |
| 2 | 78.5 tok/s | 39.2 tok/s | 75.2 ms | 145.3 ms | 20 | 0 |
| 4 | 143.7 tok/s | 35.9 tok/s | 167.2 ms | 367.1 ms | 20 | 0 |
| 8 | 185.4 tok/s | 23.2 tok/s | 309.9 ms | 506.7 ms | 24 | 0 |
| 16 | 639.3 tok/s | 40 tok/s | 310 ms | 585.2 ms | 48 | 0 |
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 356.3 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 64.2 tok/s | 64.2 tok/s | 46.2 ms | 65.4 ms | 20 | 0 |
| 2 | 119.9 tok/s | 59.9 tok/s | 93.1 ms | 121.5 ms | 20 | 0 |
| 4 | 110.5 tok/s | 27.6 tok/s | 173.8 ms | 303 ms | 20 | 0 |
| 8 | 356.3 tok/s | 44.5 tok/s | 295 ms | 429.8 ms | 24 | 0 |
Between 4 and 8 streams the aggregate rose 3.22× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 292.1 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 61.9 tok/s | 61.9 tok/s | 31.1 ms | 39.2 ms | 20 | 0 |
| 2 | 117.1 tok/s | 58.5 tok/s | 51.4 ms | 96.2 ms | 20 | 0 |
| 4 | 111.2 tok/s | 27.8 tok/s | 124.7 ms | 127 ms | 20 | 0 |
| 8 | 292.1 tok/s | 36.5 tok/s | 205.2 ms | 212.9 ms | 24 | 0 |
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
Gemma 3 27B Instruct · Q4_K_M — peaks at 221.7 tok/s aggregate
| Streams | Aggregate | Per stream | TTFT p50 | TTFT p95 | Samples | Failed |
|---|
| 1 | 48.6 tok/s | 48.6 tok/s | 65.6 ms | 90.5 ms | 20 | 0 |
| 2 | 90.1 tok/s | 45.1 tok/s | 94.8 ms | 180 ms | 20 | 0 |
| 4 | 84.3 tok/s | 21.1 tok/s | 327.8 ms | 463.9 ms | 20 | 0 |
| 8 | 221.7 tok/s | 27.7 tok/s | 375.8 ms | 449 ms | 24 | 0 |
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Concurrency was capped by VRAM on this card, not by compute.
ONE FLAGFlash attention, measured at depth
LM Studio and Ollama turn this on by default. Whether it helps, and by how much, depends on the model and only shows up once the KV cache is real — which is why it is measured at depth rather than on an empty cache.
gpt-oss 20B · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 250.7 | 4,999.9 | 11.8 GB | 3,968 |
| on | 272.1 | 8,532.1 | 11.4 GB | 3,968 |
Qwen3 30B A3B Instruct 2507 · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 197.3 | 3,146.2 | 18.4 GB | 3,968 |
| on | 218.9 | 5,947.8 | 18.3 GB | 3,968 |
Qwen3 4B Instruct 2507 · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 204.9 | 4,501.9 | 3.8 GB | 3,968 |
| on | 220 | 9,815.4 | 3.7 GB | 3,968 |
GLM 4.7 Flash · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 95.2 | 2,856.5 | 17.9 GB | 3,968 |
| on | 167.3 | 3,970.8 | 17.9 GB | 3,968 |
Llama 3.1 8B Instruct · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 148 | 4,432.8 | 5.7 GB | 3,968 |
| on | 155.1 | 8,314.7 | 5.5 GB | 3,968 |
Qwen3 8B · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 140.2 | 4,033.9 | 5.8 GB | 3,968 |
| on | 147.3 | 7,762.8 | 5.7 GB | 3,968 |
Qwen3.5 9B · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 138.8 | 6,382.9 | 5.9 GB | 3,968 |
| on | 140.7 | 7,066.9 | 5.9 GB | 3,968 |
Gemma 3 12B Instruct · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 92.9 | 4,623.6 | 8.5 GB | 3,968 |
| on | 95.3 | 6,008.2 | 8.5 GB | 3,968 |
Qwen3 14B · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 85.3 | 2,726.9 | 9.6 GB | 3,968 |
| on | 88.8 | 4,645.8 | 9.4 GB | 3,968 |
Mistral Small 3.2 24B Instruct · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 58.4 | 2,424.5 | 14.5 GB | 3,968 |
| on | 59.7 | 3,535.2 | 14.4 GB | 3,968 |
Gemma 3 27B Instruct · Q4_K_M
| -fa | Decode tok/s | Prefill tok/s | Peak VRAM | Depth |
|---|
| off | 46.4 | 2,330.7 | 17.4 GB | 3,968 |
| on | 47.2 | 2,973.3 | 17.3 GB | 3,968 |
WATTS AND MONEYWhat a million tokens costs
Power sampled at the card during generation, not the TDP on the spec sheet. Cost per million tokens uses the rate this pod was billed at; on your own hardware the same energy figures apply against your electricity price.
| Model | Quant | Avg W | Peak W | Peak °C | tok/W | kWh/Mtok | $/Mtok single | $/Mtok batched |
|---|
| gpt-oss 20B | Q4_K_M | 210 | 338 | 47 | 1.38 | 0.202 | $0.66 | $0.22 |
| Qwen3 30B A3B Instruct 2507 | Q4_K_M | 181 | 307 | 45 | 1.43 | 0.195 | $0.74 | $0.24 |
| Qwen3 4B Instruct 2507 | Q4_K_M | 268 | 332 | 49 | 0.97 | 0.288 | $0.74 | $0.15 |
| Qwen3 8B (Q3_K_M) | Q3_K_M | 338 | 414 | 55 | 0.55 | 0.504 | $1.03 | $0.17 |
| GLM 4.7 Flash | Q4_K_M | 187 | 302 | 44 | 1 | 0.278 | $1.03 | $0.23 |
| Llama 3.1 8B Instruct | Q4_K_M | 299 | 361 | 49 | 0.57 | 0.486 | $1.12 | $0.16 |
| Qwen3 8B | Q4_K_M | 298 | 356 | 51 | 0.55 | 0.506 | $1.17 | $0.18 |
| Qwen3 8B (Q5_K_M) | Q5_K_M | 298 | 359 | 52 | 0.49 | 0.572 | $1.32 | $0.19 |
| Qwen3.5 9B | Q4_K_M | 291 | 357 | 52 | 0.5 | 0.562 | $1.33 | $0.27 |
| Qwen3 8B (Q6_K) | Q6_K | 326 | 390 | 57 | 0.39 | 0.707 | $1.49 | $0.2 |
| Qwen3 8B (Q8_0) | Q8_0 | 267 | 314 | 49 | 0.39 | 0.715 | $1.85 | $0.21 |
| Gemma 3 12B Instruct | Q4_K_M | 302 | 366 | 51 | 0.34 | 0.823 | $1.88 | $0.29 |
| Qwen3 14B | Q4_K_M | 316 | 379 | 53 | 0.3 | 0.915 | $2 | $0.26 |
| Gemma 3 12B Instruct (Q5_K_M) | Q5_K_M | 305 | 362 | 53 | 0.3 | 0.932 | $2.11 | $0.3 |
| Gemma 3 12B Instruct (Q8_0) | Q8_0 | 266 | 329 | 53 | 0.25 | 1.129 | $2.92 | $0.54 |
| Mistral Small 3.2 24B Instruct | Q4_K_M | 329 | 413 | 59 | 0.19 | 1.461 | $3.07 | $0.66 |
| Gemma 3 27B Instruct | Q4_K_M | 323 | 409 | 59 | 0.15 | 1.815 | $3.88 | $0.86 |
THE MACHINEWhat this ran on, in full
A measurement is a claim about a machine, and a machine nobody described is a claim nobody can check. Fields that need root — the DMI table, which a container does not have — say so rather than being dropped: a missing row and an unreadable one mean different things.
| Property | Value |
|---|
| CPU | AMD EPYC 7452 32-Core Processor |
| Cores | 32 physical / 64 logical |
| Cores this process could use | 10 |
| Threads given to llama-bench | 10 |
| NUMA nodes | 1 |
| RAM | 504 GB |
| Kernel | 6.8.0-124-generic |
| OS | Ubuntu 24.04.1 LTS |
| System | To Be Filled By O.E.M. ROMED8-2T/BCM |
| BIOS | P4.10 06/05/2025 |
| Containerised | yes |
| Disks | nvme0n1 Samsung SSD 980 PRO 250GB 250 GB, nvme1n1 SAMSUNG MZQL27T6HBLA-00A07 7682 GB |
| GPU | VRAM | PCIe link | Power limit | Max SM / mem clock |
|---|
| 0: NVIDIA GeForce RTX 4090 | 24.0 GB | gen 1 ×16 (of gen 4 ×16) | 450 W | 3,135 / 10,501 MHz |
▸ WANT ONE FOR YOUR MACHINE?The NVIDIA GeForce RTX 4090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.
You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.
ASK ABOUT A REPORT →fitmyllm@gmail.com
HOW THIS WAS PRODUCEDA pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-4090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →