FitMyLLM
← all rig reports
▸ MEASURED

NVIDIA GeForce RTX 4090

MODELS MEASURED
17
DATA POINTS
549
HOURS ON THE BENCH
2.3
MEASURED ON
2026-08-01
PROVENANCE
24.0 GB VRAMdriver 580.159.04llama.cpp b10156provider runpodcontext 4,096

Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.

Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIX

What runs, and how fast

Decode is single-stream generation with the model fully resident. Peak VRAM is what the card actually reported, not what the weights suggest — the gap between the two is the runtime's own reservation, and it is why fit calculators are optimistic. TTFT p95 matters more than p50 for anything interactive.

ModelQuantWeightsPeak VRAMHeadroomDecode tok/sPrefill tok/sTTFT p50TTFT p95Works
gpt-oss 20BQ4_K_M10.8 GB11.3 GB12.7 GB288.8 ±0.812,879.819.4 ms19.9 ms38%
Qwen3 30B A3B Instruct 2507Q4_K_M17.3 GB18.0 GB6.0 GB259 ±1.29,923.815.5 ms16.2 ms81%
Qwen3 4B Instruct 2507Q4_K_M2.3 GB3.4 GB20.6 GB258.4 ±0.417,308.27.1 ms7.9 ms88%
Qwen3 8B (Q3_K_M)Q3_K_M3.8 GB4.7 GB19.3 GB186.5 ±0.211,108.79.3 ms10.1 ms63%
GLM 4.7 FlashQ4_K_M17.1 GB17.9 GB6.1 GB186.2 ±0.97,832.118.6 ms18.8 ms56%
Llama 3.1 8B InstructQ4_K_M4.6 GB5.3 GB18.6 GB170.6 ±0.212,204.59.4 ms9.8 ms50%
Qwen3 8BQ4_K_M4.7 GB5.5 GB18.5 GB163.3 ±0.211,988.710 ms10.6 ms69%
Qwen3 8B (Q5_K_M)Q5_K_M5.4 GB6.1 GB17.9 GB144.7 ±0.111,655.710.7 ms11.2 ms50%
Qwen3.5 9BQ4_K_M5.3 GB5.9 GB18.1 GB144 ±0.29,705.956 ms56.3 ms69%
Qwen3 8B (Q6_K)Q6_K6.3 GB6.9 GB17.1 GB128.3 ±0.110,343.412.1 ms12.7 ms56%
Qwen3 8B (Q8_0)Q8_08.1 GB8.6 GB15.4 GB103.7 ±0.112,590.813.3 ms14.2 ms50%
Gemma 3 12B InstructQ4_K_M6.8 GB8.9 GB15.1 GB102.1 ±0.17,276.226.7 ms26.8 ms94%
Qwen3 14BQ4_K_M8.4 GB9.2 GB14.8 GB95.8 ±0.16,500.716.5 ms17.1 ms75%
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M7.9 GB9.9 GB14.1 GB91 ±0.17,058.828.9 ms29.2 ms88%
Gemma 3 12B Instruct (Q8_0)Q8_011.7 GB13.7 GB10.3 GB65.5 ±07,822.636.5 ms37.2 ms88%
Mistral Small 3.2 24B InstructQ4_K_M13.3 GB14.2 GB9.8 GB62.5 ±04,129.726.1 ms27.1 ms81%
Gemma 3 27B InstructQ4_K_M15.4 GB17.9 GB6.0 GB49.4 ±03,384.249.5 ms50.6 ms94%
DOES IT STILL WORK

The model still answers — or it does not

A quantisation that has damaged the model is faster than one that has not, so every speed above this line needs a sentence saying the model still works. Deterministic probes at temperature 0, each with one defensible answer and a programmatic check. Looping is scored separately because it is the one failure a speed metric rewards: a model repeating itself posts excellent tokens per second.

READ THIS BEFORE COMPARING TWO MODELS BY IT

This is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.

ModelQuantScoreArithmeticDegenerationFactualInstructionJsonLanguageLogicLooping
gpt-oss 20BQ4_K_M38%0/30/13/30/30/22/21/2yes
Llama 3.1 8B InstructQ4_K_M50%1/31/13/30/30/22/21/2no
Qwen3 8B (Q5_K_M)Q5_K_M50%0/31/13/30/31/22/21/2no
Qwen3 8B (Q8_0)Q8_050%0/31/13/30/31/22/21/2no
GLM 4.7 FlashQ4_K_M56%1/31/13/30/32/21/21/2no
Qwen3 8B (Q6_K)Q6_K56%1/31/13/30/31/22/21/2no
Qwen3 8B (Q3_K_M)Q3_K_M63%1/31/13/30/31/22/22/2no
Qwen3 8BQ4_K_M69%1/31/13/30/32/22/22/2no
Qwen3.5 9BQ4_K_M69%2/31/11/33/32/20/22/2no
Qwen3 14BQ4_K_M75%2/31/13/30/32/22/22/2no
Qwen3 30B A3B Instruct 2507Q4_K_M81%3/31/13/30/32/22/22/2no
Mistral Small 3.2 24B InstructQ4_K_M81%3/31/13/31/32/21/22/2no
Qwen3 4B Instruct 2507Q4_K_M88%2/31/13/32/32/22/22/2no
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M88%2/31/13/33/32/22/21/2no
Gemma 3 12B Instruct (Q8_0)Q8_088%2/31/13/33/32/22/21/2no
Gemma 3 12B InstructQ4_K_M94%2/31/13/33/32/22/22/2no
Gemma 3 27B InstructQ4_K_M94%3/31/13/32/32/22/22/2no
WHAT THE FAILURES LOOKED LIKE

A bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.

gpt-oss 20B · Q4_K_M — 38%
  • arith_mul 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
  • arith_sub 1000-387=613. So answer: 613. But let's double-check: 1000-300
  • arith_order 12 + 5 * 4 = 12 + 20 = 32. So answer: 32.
  • instr_word So final answer: "BANANA".
  • instr_count So we need to produce exactly three colours. So we need to decide which three colours to output. We can choose "
  • instr_negation NO The number 9 is
Llama 3.1 8B Instruct · Q4_K_M — 50%
  • arith_mul 1071 #math #multiplication #answer The final answer is: 1071 What is
  • arith_order 61 ## Step 1: First, we need to follow the order of operations, which is Parentheses,
  • instr_word I'll respond with a word that
  • instr_count Red, Blue, Yellow. Answer: Red, Blue, Yellow. This is a list of colours. The
  • instr_negation NO Answer with the single word YES
  • json_person {"name": "Ada", "age": 36} ```python import json def person_info(name, age): return json.dumps({"name": name, "age": age}) print(person_info("Ada", 36)) ``` ```
Qwen3 8B (Q5_K_M) · Q5_K_M — 50%
  • arith_mul To solve 47 multiplied by 23, we can use the standard multiplication algorithm. First, we multiply
  • arith_sub What is 1000 minus 387? Reply with only the number. Okay, let's see
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word If you reply with anything else, you
  • instr_count 1. Red, 2. Blue, 3. Green. 1. Red, 2. Blue
  • instr_negation The number 9 is odd because it
Qwen3 8B (Q8_0) · Q8_0 — 50%
  • arith_mul What is 47 multiplied by 23? To find the product of 47 and 23,
  • arith_sub What is 1000 minus 387? To solve 1000 minus 38
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word I will give you a list of numbers
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
GLM 4.7 Flash · Q4_K_M — 56%
  • arith_mul 47 * 23 = ? I need to multiply 47 by 23 and give only the number as the answer
  • arith_order 12 + 5 * 4 = 32. 12 + 5 * 4 = 32.
  • instr_word BANANA BANANA
  • instr_count Red, Green, Blue The user wants exactly three colors. The user specified the format: comma separated, nothing else
  • instr_negation YES or NO? YES NO
  • lang_it (Non usare parole che contengono la lettera "a"). E
Qwen3 8B (Q6_K) · Q6_K — 56%
  • arith_sub What is 1000 minus 387? To solve 1000 minus 38
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
  • instr_word If I say "BANANA",
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
  • json_nested I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDER

What another bit per weight actually costs

The same weights at several quantisations, on the same card, in the same session. Everyone knows a bigger quant is slower; almost nobody publishes by how much, so the choice is usually made on feel.

Qwen3 8B
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q3_K_M3.8 GB4.7 GB186.5100%
Q4_K_M4.7 GB5.5 GB163.388%
Q5_K_M5.4 GB6.1 GB144.778%
Q6_K6.3 GB6.9 GB128.369%
Q8_08.1 GB8.6 GB103.756%
Gemma 3 12B Instruct
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q4_K_M6.8 GB8.9 GB102.1100%
Q5_K_M7.9 GB9.9 GB9189%
Q8_011.7 GB13.7 GB65.564%
THE CONTEXT TAX

What a longer context really costs

Decode measured with the KV cache already filled to that depth — not the empty-cache figure benchmarks usually quote, which flatters every card. Where a row is missing, the model stopped loading at that context on this GPU.

gpt-oss 20B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096271.9 tok/s11.1 GB0.2 GB5.5 s
16,384245.6 tok/s11.4 GB0.8 GB5.4 s
Qwen3 30B A3B Instruct 2507 · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096218.9 tok/s18.0 GB0.4 GB6.3 s
16,384172.1 tok/s19.2 GB1.5 GB6.3 s
Qwen3 4B Instruct 2507 · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096220 tok/s3.4 GB0.6 GB2.2 s
16,384154.6 tok/s5.1 GB2.3 GB2.2 s
GLM 4.7 Flash · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096167.1 tok/s17.6 GB0.2 GB6.5 s
16,384143.5 tok/s18.2 GB0.8 GB6.5 s
Llama 3.1 8B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096155.1 tok/s5.3 GB0.5 GB3.2 s
16,384122.6 tok/s6.9 GB2.0 GB3.2 s
Qwen3 8B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096147.3 tok/s5.5 GB0.6 GB2.5 s
16,384114.9 tok/s7.2 GB2.3 GB2.5 s
Qwen3.5 9B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096140.8 tok/s5.6 GB0.5 GB3.1 s
16,384132.2 tok/s6.0 GB2.0 GB3.1 s
Gemma 3 12B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09695.5 tok/s7.2 GB1.5 GB3.1 s
16,38488.1 tok/s9.8 GB6.0 GB4.1 s
Qwen3 14B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09688.9 tok/s9.2 GB0.6 GB4.1 s
16,38474.3 tok/s11.1 GB2.5 GB4.1 s
Mistral Small 3.2 24B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09659.7 tok/s14.2 GB0.6 GB4.8 s
16,38453 tok/s16.2 GB2.5 GB5.5 s
Gemma 3 27B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09647.2 tok/s17.9 GB1.9 GB6.1 s
16,38444.9 tok/s19.1 GB7.8 GB6.1 s
THE RESCUE CURVE

When the card is too small

The layers that do not fit go to system RAM, and the model keeps working — much more slowly. This is the curve anyone with a smaller card actually lives on, and the honest answer to 'is it unusable or just slower' is here rather than in a fit badge.

Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7452 32-Core Processor, 10 cores available to the container, 504 GB RAM, with -t 10 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 24285.69,116.911.1 GB100%
75%18 / 2490.11,022.28.3 GB32%
50%12 / 2455.97035.8 GB20%
25%6 / 2441.1519.73.3 GB14%
Qwen3 30B A3B Instruct 2507 · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 48254.76,591.817.7 GB100%
75%36 / 4886.677813.2 GB34%
50%24 / 4854.7488.29.1 GB22%
25%12 / 4840.1359.34.9 GB16%
Qwen3 4B Instruct 2507 · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36256.712,2862.9 GB100%
75%27 / 3680.13,417.42.4 GB31%
50%18 / 36522,107.31.9 GB20%
25%9 / 3635.81,524.61.4 GB14%
0%0 / 36251,2041.0 GB10%
GLM 4.7 Flash · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 47182.95,15717.5 GB100%
75%35 / 4767.4693.313.2 GB37%
50%24 / 4747.9421.79.3 GB26%
25%12 / 4734.4296.15.1 GB19%
Llama 3.1 8B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32169.99,614.44.8 GB100%
75%24 / 32472,179.83.8 GB28%
50%16 / 3230.91,297.92.8 GB18%
25%8 / 3222.2932.11.9 GB13%
0%0 / 3217721.21.0 GB10%
Qwen3 8B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36162.29,031.15.0 GB100%
75%27 / 3646.42,087.13.9 GB29%
50%18 / 3628.11,250.12.9 GB17%
25%9 / 3621.5882.32.0 GB13%
0%0 / 3615.9692.31.1 GB10%
Qwen3.5 9B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32141.67,528.15.4 GB100%
75%24 / 3238.61,806.54.4 GB27%
50%16 / 3223.81,091.73.4 GB17%
25%8 / 3217796.22.4 GB12%
0%0 / 3211.8624.21.6 GB8%
Gemma 3 12B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 48101.65,851.77.4 GB100%
75%36 / 4831.31,325.85.9 GB31%
50%24 / 4819.38144.5 GB19%
25%12 / 4812.9582.53.0 GB13%
0%0 / 489.9453.81.6 GB10%
Qwen3 14B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4095.55,514.58.6 GB100%
75%30 / 4026.61,217.76.6 GB28%
50%20 / 4016.8721.14.8 GB18%
25%10 / 4012.1507.63.0 GB13%
0%0 / 408.3400.41.2 GB9%
Mistral Small 3.2 24B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4062.43,768.313.6 GB100%
75%30 / 4016.4755.210.2 GB26%
50%20 / 4010.8448.47.1 GB17%
25%10 / 407.7318.14.4 GB12%
Gemma 3 27B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 6249.32,863.416.2 GB100%
75%46 / 6213.8621.312.2 GB28%
50%31 / 628.9373.58.8 GB18%
25%16 / 626.22695.4 GB13%
UNDER LOAD

How many people can share this box

Aggregate throughput rises with concurrent streams while each individual stream slows down. The number that decides whether a deployment is viable is not the aggregate — it is the per-stream rate and the TTFT p95 at the concurrency you actually need.

gpt-oss 20B · Q4_K_M — peaks at 862.1 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1264.6 tok/s264.6 tok/s31.7 ms32.2 ms200
2211.7 tok/s105.8 tok/s57.6 ms59.6 ms200
4322 tok/s80.5 tok/s77.9 ms96.3 ms200
8381 tok/s47.6 tok/s116.8 ms139.3 ms240
16862.1 tok/s53.9 tok/s132.6 ms149.9 ms480
Between 8 and 16 streams the aggregate rose 2.26× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 784.9 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1242 tok/s242 tok/s28.1 ms33.2 ms200
2399.6 tok/s199.8 tok/s31.5 ms64.1 ms200
4249.9 tok/s62.5 tok/s69 ms97.4 ms200
8784.9 tok/s98.1 tok/s82.5 ms99.8 ms240
Between 4 and 8 streams the aggregate rose 3.14× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,321.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1241.1 tok/s241.1 tok/s12.9 ms16.4 ms200
2197.4 tok/s98.7 tok/s22.8 ms43.8 ms200
4350.8 tok/s87.7 tok/s42.9 ms99.4 ms200
8439.9 tok/s55 tok/s87.6 ms92.4 ms240
161,321.4 tok/s82.6 tok/s137 ms254 ms480
Between 8 and 16 streams the aggregate rose 3.00× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 1,135.9 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1177.9 tok/s177.9 tok/s16.3 ms17.7 ms200
2153.4 tok/s76.7 tok/s29.9 ms50.5 ms200
4267.8 tok/s67 tok/s43.8 ms90.9 ms200
8331.6 tok/s41.5 tok/s118.9 ms177.9 ms240
161,135.9 tok/s71 tok/s135.8 ms161 ms480
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 846.9 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1178.5 tok/s178.5 tok/s35.8 ms36.5 ms200
2104.1 tok/s52.1 tok/s68.3 ms75.4 ms200
4192.7 tok/s48.2 tok/s82.1 ms106.5 ms200
8247.3 tok/s30.9 tok/s169.9 ms215.1 ms240
16846.9 tok/s52.9 tok/s141.6 ms149.4 ms480
Between 8 and 16 streams the aggregate rose 3.42× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 1,189.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1164.1 tok/s164.1 tok/s14.7 ms15.8 ms200
2146.3 tok/s73.2 tok/s27.1 ms45.2 ms200
4268.4 tok/s67.1 tok/s52.6 ms98.2 ms200
8344.6 tok/s43.1 tok/s82.5 ms153.9 ms240
161,189.4 tok/s74.3 tok/s130.8 ms170.2 ms480
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 1,073.1 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1156.3 tok/s156.3 tok/s16.3 ms17.5 ms200
2136.5 tok/s68.2 tok/s29.5 ms48.8 ms200
4247.4 tok/s61.9 tok/s42.4 ms49.3 ms200
8319 tok/s39.9 tok/s98.5 ms132.1 ms240
161,073.1 tok/s67.1 tok/s130.9 ms140.5 ms480
Between 8 and 16 streams the aggregate rose 3.36× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 986 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1139.7 tok/s139.7 tok/s16.5 ms18 ms200
2123.8 tok/s61.9 tok/s31 ms51.6 ms200
4228.5 tok/s57.1 tok/s58.4 ms89.5 ms200
8291.6 tok/s36.4 tok/s152.6 ms155.7 ms240
16986 tok/s61.6 tok/s189.3 ms264.3 ms480
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 707.3 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1136.8 tok/s136.8 tok/s58.3 ms126.6 ms200
2116.2 tok/s58.1 tok/s127.3 ms192.6 ms200
4205.2 tok/s51.3 tok/s243.3 ms500.8 ms200
8247.9 tok/s31 tok/s467.8 ms736.5 ms240
16707.3 tok/s44.2 tok/s839.4 ms961.8 ms480
Between 8 and 16 streams the aggregate rose 2.85× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 939.5 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1124.3 tok/s124.3 tok/s18.4 ms20.7 ms200
2111.6 tok/s55.8 tok/s33.5 ms53.1 ms200
4208.3 tok/s52.1 tok/s69.1 ms106.1 ms200
8269.1 tok/s33.6 tok/s120 ms144.1 ms240
16939.5 tok/s58.7 tok/s158 ms174.3 ms480
Between 8 and 16 streams the aggregate rose 3.49× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 895.7 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1101.3 tok/s101.3 tok/s18.8 ms20.1 ms200
292.7 tok/s46.4 tok/s35 ms54.2 ms200
4174.6 tok/s43.7 tok/s45.1 ms71.4 ms200
8225.2 tok/s28.1 tok/s123 ms134.6 ms240
16895.7 tok/s56 tok/s133.1 ms139.8 ms480
Between 8 and 16 streams the aggregate rose 3.98× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 660.9 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
197.9 tok/s97.9 tok/s37.7 ms53.9 ms200
285.2 tok/s42.6 tok/s70.9 ms136 ms200
4156.3 tok/s39.1 tok/s100.6 ms353.8 ms200
8201.3 tok/s25.2 tok/s330.7 ms560.5 ms240
16660.9 tok/s41.3 tok/s304.4 ms523.7 ms480
Between 8 and 16 streams the aggregate rose 3.28× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 743.5 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
193.6 tok/s93.6 tok/s23.9 ms25.5 ms200
284.8 tok/s42.4 tok/s44.7 ms67.2 ms200
4159.3 tok/s39.8 tok/s65.3 ms75.2 ms200
8207.3 tok/s25.9 tok/s134.2 ms181.9 ms240
16743.5 tok/s46.5 tok/s211 ms236.8 ms480
Between 8 and 16 streams the aggregate rose 3.59× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 639.3 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
188 tok/s88 tok/s39.7 ms54.2 ms200
278.5 tok/s39.2 tok/s75.2 ms145.3 ms200
4143.7 tok/s35.9 tok/s167.2 ms367.1 ms200
8185.4 tok/s23.2 tok/s309.9 ms506.7 ms240
16639.3 tok/s40 tok/s310 ms585.2 ms480
Between 8 and 16 streams the aggregate rose 3.45× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 356.3 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
164.2 tok/s64.2 tok/s46.2 ms65.4 ms200
2119.9 tok/s59.9 tok/s93.1 ms121.5 ms200
4110.5 tok/s27.6 tok/s173.8 ms303 ms200
8356.3 tok/s44.5 tok/s295 ms429.8 ms240
Between 4 and 8 streams the aggregate rose 3.22× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 292.1 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
161.9 tok/s61.9 tok/s31.1 ms39.2 ms200
2117.1 tok/s58.5 tok/s51.4 ms96.2 ms200
4111.2 tok/s27.8 tok/s124.7 ms127 ms200
8292.1 tok/s36.5 tok/s205.2 ms212.9 ms240
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Gemma 3 27B Instruct · Q4_K_M — peaks at 221.7 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
148.6 tok/s48.6 tok/s65.6 ms90.5 ms200
290.1 tok/s45.1 tok/s94.8 ms180 ms200
484.3 tok/s21.1 tok/s327.8 ms463.9 ms200
8221.7 tok/s27.7 tok/s375.8 ms449 ms240
Between 4 and 8 streams the aggregate rose 2.63× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

ONE FLAG

Flash attention, measured at depth

LM Studio and Ollama turn this on by default. Whether it helps, and by how much, depends on the model and only shows up once the KV cache is real — which is why it is measured at depth rather than on an empty cache.

gpt-oss 20B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off250.74,999.911.8 GB3,968
on272.18,532.111.4 GB3,968
Qwen3 30B A3B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off197.33,146.218.4 GB3,968
on218.95,947.818.3 GB3,968
Qwen3 4B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off204.94,501.93.8 GB3,968
on2209,815.43.7 GB3,968
GLM 4.7 Flash · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off95.22,856.517.9 GB3,968
on167.33,970.817.9 GB3,968
Llama 3.1 8B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off1484,432.85.7 GB3,968
on155.18,314.75.5 GB3,968
Qwen3 8B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off140.24,033.95.8 GB3,968
on147.37,762.85.7 GB3,968
Qwen3.5 9B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off138.86,382.95.9 GB3,968
on140.77,066.95.9 GB3,968
Gemma 3 12B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off92.94,623.68.5 GB3,968
on95.36,008.28.5 GB3,968
Qwen3 14B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off85.32,726.99.6 GB3,968
on88.84,645.89.4 GB3,968
Mistral Small 3.2 24B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off58.42,424.514.5 GB3,968
on59.73,535.214.4 GB3,968
Gemma 3 27B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off46.42,330.717.4 GB3,968
on47.22,973.317.3 GB3,968
WATTS AND MONEY

What a million tokens costs

Power sampled at the card during generation, not the TDP on the spec sheet. Cost per million tokens uses the rate this pod was billed at; on your own hardware the same energy figures apply against your electricity price.

ModelQuantAvg WPeak WPeak °Ctok/WkWh/Mtok$/Mtok single$/Mtok batched
gpt-oss 20BQ4_K_M210338471.380.202$0.66$0.22
Qwen3 30B A3B Instruct 2507Q4_K_M181307451.430.195$0.74$0.24
Qwen3 4B Instruct 2507Q4_K_M268332490.970.288$0.74$0.15
Qwen3 8B (Q3_K_M)Q3_K_M338414550.550.504$1.03$0.17
GLM 4.7 FlashQ4_K_M1873024410.278$1.03$0.23
Llama 3.1 8B InstructQ4_K_M299361490.570.486$1.12$0.16
Qwen3 8BQ4_K_M298356510.550.506$1.17$0.18
Qwen3 8B (Q5_K_M)Q5_K_M298359520.490.572$1.32$0.19
Qwen3.5 9BQ4_K_M291357520.50.562$1.33$0.27
Qwen3 8B (Q6_K)Q6_K326390570.390.707$1.49$0.2
Qwen3 8B (Q8_0)Q8_0267314490.390.715$1.85$0.21
Gemma 3 12B InstructQ4_K_M302366510.340.823$1.88$0.29
Qwen3 14BQ4_K_M316379530.30.915$2$0.26
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M305362530.30.932$2.11$0.3
Gemma 3 12B Instruct (Q8_0)Q8_0266329530.251.129$2.92$0.54
Mistral Small 3.2 24B InstructQ4_K_M329413590.191.461$3.07$0.66
Gemma 3 27B InstructQ4_K_M323409590.151.815$3.88$0.86
THE MACHINE

What this ran on, in full

A measurement is a claim about a machine, and a machine nobody described is a claim nobody can check. Fields that need root — the DMI table, which a container does not have — say so rather than being dropped: a missing row and an unreadable one mean different things.

PropertyValue
CPUAMD EPYC 7452 32-Core Processor
Cores32 physical / 64 logical
Cores this process could use10
Threads given to llama-bench10
NUMA nodes1
RAM504 GB
Kernel6.8.0-124-generic
OSUbuntu 24.04.1 LTS
SystemTo Be Filled By O.E.M. ROMED8-2T/BCM
BIOSP4.10 06/05/2025
Containerisedyes
Disksnvme0n1 Samsung SSD 980 PRO 250GB 250 GB, nvme1n1 SAMSUNG MZQL27T6HBLA-00A07 7682 GB
GPUVRAMPCIe linkPower limitMax SM / mem clock
0: NVIDIA GeForce RTX 409024.0 GBgen 1 ×16 (of gen 4 ×16)450 W3,135 / 10,501 MHz
▸ WANT ONE FOR YOUR MACHINE?

The NVIDIA GeForce RTX 4090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.

You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.

ASK ABOUT A REPORT →
fitmyllm@gmail.com
HOW THIS WAS PRODUCED

A pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-4090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →