FitMyLLM
← all rig reports
▸ MEASURED

NVIDIA GeForce RTX 3090

MODELS MEASURED
17
DATA POINTS
551
HOURS ON THE BENCH
1.8
MEASURED ON
2026-07-31
PROVENANCE
24.0 GB VRAMdriver 580.126.20llama.cpp b10156provider runpodcontext 4,096

Every figure below came off this machine. Nothing on this page is derived from a formula: each model was downloaded, loaded and timed, and where a number is missing it is because the measurement failed, not because it was estimated.

Fastest single stream gpt-oss 20BLargest that fits Qwen3 30B A3B Instruct 2507Cheapest batched Qwen3 4B Instruct 2507
THE MATRIX

What runs, and how fast

Decode is single-stream generation with the model fully resident. Peak VRAM is what the card actually reported, not what the weights suggest — the gap between the two is the runtime's own reservation, and it is why fit calculators are optimistic. TTFT p95 matters more than p50 for anything interactive.

ModelQuantWeightsPeak VRAMHeadroomDecode tok/sPrefill tok/sTTFT p50TTFT p95Works
gpt-oss 20BQ4_K_M10.8 GB11.2 GB12.8 GB227.4 ±1.15,859.225.9 ms28.7 ms44%
Qwen3 30B A3B Instruct 2507Q4_K_M17.3 GB17.9 GB6.1 GB207.6 ±0.84,323.820.3 ms22.4 ms88%
Qwen3 4B Instruct 2507Q4_K_M2.3 GB3.3 GB20.7 GB206.4 ±0.98,258.29.4 ms14.9 ms88%
GLM 4.7 FlashQ4_K_M17.1 GB17.5 GB6.5 GB149.3 ±0.93,653.822.9 ms25.7 ms56%
Llama 3.1 8B InstructQ4_K_M4.6 GB5.2 GB18.8 GB145.3 ±0.35,189.712.8 ms13.4 ms50%
Qwen3 8BQ4_K_M4.7 GB5.3 GB18.7 GB138 ±0.25,001.213.3 ms14 ms63%
Qwen3 8B (Q5_K_M)Q5_K_M5.4 GB6.0 GB18.0 GB123.9 ±0.14,856.314 ms14.6 ms56%
Qwen3.5 9BQ4_K_M5.3 GB5.6 GB18.4 GB123.7 ±0.14,139.664.7 ms72.8 ms69%
Qwen3 8B (Q3_K_M)Q3_K_M3.8 GB4.5 GB19.5 GB119.1 ±0.14,628.213.3 ms14.6 ms63%
Qwen3 8B (Q6_K)Q6_K6.3 GB6.7 GB17.3 GB105.9 ±0.14,399.815.5 ms16.2 ms56%
Qwen3 8B (Q8_0)Q8_08.1 GB8.5 GB15.5 GB92.9 ±0.25,214.117.2 ms19.1 ms50%
Gemma 3 12B InstructQ4_K_M6.8 GB8.7 GB15.3 GB86.4 ±0.13,243.441.4 ms42 ms94%
Qwen3 14BQ4_K_M8.4 GB9.0 GB15.0 GB81.5 ±0.12,987.522.2 ms23.2 ms75%
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M7.9 GB9.8 GB14.2 GB77.8 ±03,170.844.8 ms45.8 ms88%
Gemma 3 12B Instruct (Q8_0)Q8_011.7 GB13.6 GB10.4 GB58.2 ±0.13,421.242.8 ms44 ms88%
Mistral Small 3.2 24B InstructQ4_K_M13.3 GB14.1 GB9.9 GB54 ±01,91532.3 ms33.7 ms81%
Gemma 3 27B InstructQ4_K_M15.4 GB17.8 GB6.2 GB42.6 ±01,482.684.7 ms85.2 ms94%
DOES IT STILL WORK

The model still answers — or it does not

A quantisation that has damaged the model is faster than one that has not, so every speed above this line needs a sentence saying the model still works. Deterministic probes at temperature 0, each with one defensible answer and a programmatic check. Looping is scored separately because it is the one failure a speed metric rewards: a model repeating itself posts excellent tokens per second.

READ THIS BEFORE COMPARING TWO MODELS BY IT

This is a smoke test for quantisation damage, not a quality benchmark. Fifteen probes cannot tell you which model reasons better — MMLU exists and we are not reimplementing it on a rented pod. What they catch is a quant, an offload setting or a runtime flag that has broken the model while leaving the speed column looking excellent. Read a low score as do not trust the speeds above, not as this model is bad.

ModelQuantScoreArithmeticDegenerationFactualInstructionJsonLanguageLogicLooping
gpt-oss 20BQ4_K_M44%0/31/12/30/30/22/22/2no
Llama 3.1 8B InstructQ4_K_M50%1/31/13/30/30/22/21/2no
Qwen3 8B (Q8_0)Q8_050%0/31/13/30/31/22/21/2no
GLM 4.7 FlashQ4_K_M56%1/30/13/30/32/22/21/2yes
Qwen3 8B (Q5_K_M)Q5_K_M56%1/31/13/30/31/22/21/2no
Qwen3 8B (Q6_K)Q6_K56%1/31/13/30/31/22/21/2no
Qwen3 8BQ4_K_M63%0/31/13/30/32/22/22/2no
Qwen3 8B (Q3_K_M)Q3_K_M63%1/31/13/30/31/22/22/2no
Qwen3.5 9BQ4_K_M69%2/31/11/33/32/20/22/2no
Qwen3 14BQ4_K_M75%2/31/13/30/32/22/22/2no
Mistral Small 3.2 24B InstructQ4_K_M81%3/31/13/31/32/21/22/2no
Qwen3 30B A3B Instruct 2507Q4_K_M88%3/31/13/31/32/22/22/2no
Qwen3 4B Instruct 2507Q4_K_M88%2/31/13/32/32/22/22/2no
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M88%2/31/13/33/32/22/21/2no
Gemma 3 12B Instruct (Q8_0)Q8_088%2/31/13/33/32/22/21/2no
Gemma 3 12B InstructQ4_K_M94%3/31/13/33/32/22/21/2no
Gemma 3 27B InstructQ4_K_M94%3/31/13/32/32/22/22/2no
WHAT THE FAILURES LOOKED LIKE

A bare score hides whether the model made one arithmetic slip or emitted four hundred words in the wrong language. These are its actual answers.

gpt-oss 20B · Q4_K_M — 44%
  • arith_mul 47*23=1081. So answer: 1081.<|end|><|start|>assistant<|channel|>final 1081
  • arith_sub 1000-387=613. So answer: 613. But let's double-check: 1000-387
  • arith_order 12 + 5 * 4 = 12 + 20 = 32. So answer: 32.
  • instr_word So final answer: BANANA.<|end|>
  • instr_count So we need to produce exactly three colours. So we need to decide which three colours to output. We can choose "
  • instr_negation NO The number 9 is
Llama 3.1 8B Instruct · Q4_K_M — 50%
  • arith_mul 1071. 1071. 1071. 1071. 1071. 1071.
  • arith_order 61 ## Step 1: First, we need to follow the order of operations, which is Parentheses,
  • instr_word I'll respond with a word that
  • instr_count Red, Blue, Yellow. Answer: Red, Blue, Yellow. This is a list of colours. The
  • instr_negation NO. Is the number 9 odd
  • json_person {"name": "Ada", "age": 36} ```python import json def person_info(name, age): return json.dumps({"name": name, "age": age}) print(person_info("Ada", 36)) ``` ```
Qwen3 8B (Q8_0) · Q8_0 — 50%
  • arith_mul What is 47 multiplied by 23? To find the product of 47 and 23,
  • arith_sub What is 1000 minus 387? Reply with only the number. What is 10
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word If I say "I love you",
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
GLM 4.7 Flash · Q4_K_M — 56%
  • arith_mul 47 * 23 = ? I need to multiply 47 by 23 and give only the number as the answer
  • arith_order 12 + 5 * 4 = 32. 12 + 5 * 4 = 32.
  • instr_word BANANA The user wants a
  • instr_count Red, Green, Blue The user wants exactly three colors, separated by commas, with no other text. The user
  • instr_negation YES or NO? YES NO
  • logic_compare 9.11 9.9 9.11 9.9 9.11 9.9
Qwen3 8B (Q5_K_M) · Q5_K_M — 56%
  • arith_sub What is 1000 minus 387? Reply with only the number. Okay, let's see
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. What is 1
  • instr_word If you reply with anything else, you
  • instr_count 1. Red, 2. Blue, 3. Green. 1. Red, 2. Blue
  • instr_negation The number 9 is odd because it
  • json_nested I need to make sure that the JSON is valid and properly formatted. Also, ensure that the keys are in double quotes and the boolean value is lowercase. Let me c
Qwen3 8B (Q6_K) · Q6_K — 56%
  • arith_sub What is 1000 minus 387? To solve 1000 minus 38
  • arith_order Also, what is 12 + 5 * 4? Reply with only the number. Also, what is
  • instr_word If I say "BANANA",
  • instr_count Red, Blue, Green. Wait, I need to list exactly three colors, comma separated, nothing else. Let
  • instr_negation The number 9 is odd because it
  • json_nested I need to return only the JSON, no other text. Please make sure the JSON is valid. I need to return only the JSON, no other text. Please make sure the JSON is
THE QUANT LADDER

What another bit per weight actually costs

The same weights at several quantisations, on the same card, in the same session. Everyone knows a bigger quant is slower; almost nobody publishes by how much, so the choice is usually made on feel.

Qwen3 8B
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q4_K_M4.7 GB5.3 GB138100%
Q5_K_M5.4 GB6.0 GB123.990%
Q3_K_M3.8 GB4.5 GB119.186%
Q6_K6.3 GB6.7 GB105.977%
Q8_08.1 GB8.5 GB92.967%
Gemma 3 12B Instruct
QuantWeightsPeak VRAMDecode tok/svs fastest rung
Q4_K_M6.8 GB8.7 GB86.4100%
Q5_K_M7.9 GB9.8 GB77.890%
Q8_011.7 GB13.6 GB58.267%
THE CONTEXT TAX

What a longer context really costs

Decode measured with the KV cache already filled to that depth — not the empty-cache figure benchmarks usually quote, which flatters every card. Where a row is missing, the model stopped loading at that context on this GPU.

gpt-oss 20B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096216.5 tok/s11.0 GB0.2 GB6.1 s
16,384203.1 tok/s11.2 GB0.8 GB6 s
Qwen3 30B A3B Instruct 2507 · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096185.2 tok/s17.9 GB0.4 GB6.8 s
16,384146 tok/s19.0 GB1.5 GB7 s
Qwen3 4B Instruct 2507 · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096177.7 tok/s3.3 GB0.6 GB2 s
16,384131.5 tok/s5.0 GB2.3 GB2 s
GLM 4.7 Flash · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096130 tok/s17.5 GB0.2 GB7.3 s
16,384107.1 tok/s18.1 GB0.8 GB7 s
Llama 3.1 8B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096132.7 tok/s5.2 GB0.5 GB3 s
16,384107.6 tok/s6.7 GB2.0 GB3 s
Qwen3 8B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096125.5 tok/s5.3 GB0.6 GB3 s
16,384100.1 tok/s7.0 GB2.3 GB3 s
Qwen3.5 9B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,096121.3 tok/s5.5 GB0.5 GB3.3 s
16,384115.3 tok/s5.8 GB2.0 GB3.3 s
Gemma 3 12B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09681.1 tok/s8.7 GB1.5 GB4.8 s
16,38475.5 tok/s9.6 GB6.0 GB4.2 s
Qwen3 14B · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09676.7 tok/s9.0 GB0.6 GB4.2 s
16,38465.3 tok/s10.9 GB2.5 GB4.3 s
Mistral Small 3.2 24B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09651.6 tok/s14.1 GB0.6 GB4.9 s
16,38446.4 tok/s16.0 GB2.5 GB4.9 s
Gemma 3 27B Instruct · Q4_K_M
ContextDecode at depthPeak VRAMKV cache (theory)Load time
4,09640.4 tok/s17.8 GB1.9 GB6.5 s
16,38439.1 tok/s19.0 GB7.8 GB7.8 s
THE RESCUE CURVE

When the card is too small

The layers that do not fit go to system RAM, and the model keeps working — much more slowly. This is the curve anyone with a smaller card actually lives on, and the honest answer to 'is it unusable or just slower' is here rather than in a fit badge.

Everything below the fully-resident row is partly a measurement of the CPU: AMD EPYC 7H12 64-Core Processor, 27 cores available to the container, 1008 GB RAM, with -t 27 passed explicitly. Your own curve moves with your CPU and your memory bandwidth, so read the shape rather than the absolute tok/s.
gpt-oss 20B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 24227.24,347.911.0 GB100%
75%18 / 2483.8811.38.1 GB37%
50%12 / 2452.5558.95.7 GB23%
25%6 / 2441.4421.73.2 GB18%
Qwen3 30B A3B Instruct 2507 · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 48205.53,094.217.6 GB100%
75%36 / 4856.6500.513.0 GB28%
50%24 / 4834.2332.18.9 GB17%
25%12 / 4824.2242.54.8 GB12%
Qwen3 4B Instruct 2507 · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36203.36,439.42.8 GB100%
75%27 / 3688.62,847.82.2 GB44%
50%18 / 36551,418.51.7 GB27%
25%9 / 3641.71,350.71.3 GB21%
0%0 / 3630.51,081.70.8 GB15%
GLM 4.7 Flash · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 47146.22,498.217.4 GB100%
75%35 / 4766.860013.0 GB46%
50%24 / 4742.8387.39.1 GB29%
25%12 / 4729.1277.74.9 GB20%
Llama 3.1 8B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32145.24,617.74.7 GB100%
75%24 / 3245.71,470.83.7 GB32%
50%16 / 3228.9986.72.7 GB20%
25%8 / 3219.57081.7 GB13%
0%0 / 3215.7560.50.9 GB11%
Qwen3 8B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 36137.84,315.54.8 GB100%
75%27 / 3647.61,690.93.8 GB35%
50%18 / 3631.91,077.62.8 GB23%
25%9 / 3622.5798.31.9 GB16%
0%0 / 3617.5651.41.0 GB13%
Qwen3.5 9B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 32122.23,534.65.3 GB100%
75%24 / 3239.51,422.24.2 GB32%
50%16 / 3225.2931.23.2 GB21%
25%8 / 3217.7694.22.3 GB14%
0%0 / 3213.15671.4 GB11%
Gemma 3 12B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4886.22,849.77.4 GB100%
75%36 / 4833.11,068.65.8 GB39%
50%24 / 4821.6692.24.3 GB25%
25%12 / 4814.8501.52.8 GB17%
0%0 / 4811.2387.31.5 GB13%
Qwen3 14B · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 4081.32,7358.5 GB100%
75%30 / 4031.2984.86.4 GB38%
50%20 / 4019.46304.6 GB24%
25%10 / 4014.8460.42.8 GB18%
0%0 / 4011.2370.41.1 GB14%
Mistral Small 3.2 24B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 40541,804.413.5 GB100%
75%30 / 4017.5633.410.0 GB32%
50%20 / 4012.1395.57.0 GB22%
25%10 / 409.4289.53.9 GB17%
Gemma 3 27B Instruct · Q4_K_M
On GPULayersDecode tok/sPrefill tok/sPeak VRAMvs fully resident
100%999 / 6242.51,388.516.1 GB100%
75%46 / 6214.3489.912.1 GB34%
50%31 / 626.8270.38.7 GB16%
25%16 / 625.8225.95.3 GB14%
UNDER LOAD

How many people can share this box

Aggregate throughput rises with concurrent streams while each individual stream slows down. The number that decides whether a deployment is viable is not the aggregate — it is the per-stream rate and the TTFT p95 at the concurrency you actually need.

gpt-oss 20B · Q4_K_M — peaks at 594 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1210.4 tok/s210.4 tok/s49.8 ms55.2 ms200
2178.3 tok/s89.1 tok/s95.3 ms98.7 ms200
4261.1 tok/s65.3 tok/s128.9 ms170.3 ms200
8290.4 tok/s36.3 tok/s230.6 ms244.8 ms240
16594 tok/s37.1 tok/s258.3 ms324.7 ms480
Qwen3 30B A3B Instruct 2507 · Q4_K_M — peaks at 761.2 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1196.8 tok/s196.8 tok/s49.4 ms53.7 ms200
2129.7 tok/s64.9 tok/s98.1 ms114.3 ms200
4226.7 tok/s56.7 tok/s123.1 ms185.8 ms200
8269.3 tok/s33.7 tok/s267 ms355.9 ms240
16761.2 tok/s47.6 tok/s258.3 ms306.7 ms480
Between 8 and 16 streams the aggregate rose 2.83× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 4B Instruct 2507 · Q4_K_M — peaks at 1,108.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1195.6 tok/s195.6 tok/s19.4 ms31.8 ms200
2164.2 tok/s82.1 tok/s34.9 ms57.4 ms200
4276.6 tok/s69.1 tok/s67.3 ms108.7 ms200
8327.6 tok/s41 tok/s182.9 ms191.1 ms240
161,108.4 tok/s69.3 tok/s253.9 ms398.6 ms480
Between 8 and 16 streams the aggregate rose 3.38× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
GLM 4.7 Flash · Q4_K_M — peaks at 566.2 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1144 tok/s144 tok/s63.9 ms64.8 ms200
2100.8 tok/s50.4 tok/s126.3 ms134.4 ms200
4175.1 tok/s43.8 tok/s155.7 ms179.5 ms200
8208.3 tok/s26 tok/s304.1 ms377.2 ms240
16566.2 tok/s35.4 tok/s326 ms371.1 ms480
Between 8 and 16 streams the aggregate rose 2.72× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Llama 3.1 8B Instruct · Q4_K_M — peaks at 923.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1140.6 tok/s140.6 tok/s26.7 ms28.4 ms200
2126.3 tok/s63.2 tok/s47.6 ms65.9 ms200
4211.9 tok/s53 tok/s95.8 ms118.5 ms200
8248 tok/s31 tok/s173.9 ms182.2 ms240
16923.4 tok/s57.7 tok/s307 ms410.3 ms480
Between 8 and 16 streams the aggregate rose 3.72× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B · Q4_K_M — peaks at 853.5 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1133.1 tok/s133.1 tok/s28.8 ms30.7 ms200
2117.2 tok/s58.6 tok/s51 ms76.6 ms200
4197.4 tok/s49.4 tok/s84.9 ms110.3 ms200
8231.9 tok/s29 tok/s206.7 ms214.1 ms240
16853.5 tok/s53.3 tok/s300.1 ms491.5 ms480
Between 8 and 16 streams the aggregate rose 3.68× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q5_K_M) · Q5_K_M — peaks at 836 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1120.1 tok/s120.1 tok/s29.3 ms31 ms200
2106.6 tok/s53.3 tok/s53.7 ms84.9 ms200
4182.7 tok/s45.7 tok/s84.3 ms161.7 ms200
8215.8 tok/s27 tok/s225.9 ms265.3 ms240
16836 tok/s52.3 tok/s393.8 ms588.2 ms480
Between 8 and 16 streams the aggregate rose 3.87× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3.5 9B · Q4_K_M — peaks at 523.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1119.2 tok/s119.2 tok/s73.9 ms158.3 ms200
2102.8 tok/s51.4 tok/s161.8 ms244.7 ms200
4165.9 tok/s41.5 tok/s380.4 ms608.2 ms200
8194.3 tok/s24.3 tok/s817.2 ms971.5 ms240
16523.4 tok/s32.7 tok/s1,327.4 ms2,100.9 ms480
Between 8 and 16 streams the aggregate rose 2.69× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q3_K_M) · Q3_K_M — peaks at 895 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1116.2 tok/s116.2 tok/s29.5 ms31.9 ms200
2104.9 tok/s52.5 tok/s54.3 ms77.5 ms200
4176.1 tok/s44 tok/s86.5 ms133.5 ms200
8211.2 tok/s26.4 tok/s179.4 ms242.5 ms240
16895 tok/s55.9 tok/s293.5 ms334 ms480
Between 8 and 16 streams the aggregate rose 4.24× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q6_K) · Q6_K — peaks at 769.8 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
1103.6 tok/s103.6 tok/s31.6 ms33.1 ms200
293.9 tok/s46.9 tok/s59.2 ms87.3 ms200
4169.1 tok/s42.3 tok/s108.5 ms131.5 ms200
8206.4 tok/s25.8 tok/s221 ms238.7 ms240
16769.8 tok/s48.1 tok/s325.6 ms364.6 ms480
Between 8 and 16 streams the aggregate rose 3.73× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 8B (Q8_0) · Q8_0 — peaks at 713.2 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
190.8 tok/s90.8 tok/s26.9 ms28.8 ms200
282.7 tok/s41.4 tok/s52.5 ms76 ms200
4155 tok/s38.8 tok/s81.9 ms166.3 ms200
8199.2 tok/s24.9 tok/s234.4 ms296.4 ms240
16713.2 tok/s44.6 tok/s313.6 ms372.8 ms480
Between 8 and 16 streams the aggregate rose 3.58× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct · Q4_K_M — peaks at 556.4 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
184 tok/s84 tok/s59.3 ms79.7 ms200
274 tok/s37 tok/s122.7 ms216.7 ms200
4122.4 tok/s30.6 tok/s179.1 ms415.3 ms200
8144.8 tok/s18.1 tok/s492.5 ms669.1 ms240
16556.4 tok/s34.8 tok/s758.6 ms876 ms480
Between 8 and 16 streams the aggregate rose 3.84× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Qwen3 14B · Q4_K_M — peaks at 613.3 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
179.7 tok/s79.7 tok/s43 ms45.4 ms200
272.9 tok/s36.5 tok/s90.5 ms105.9 ms200
4121.8 tok/s30.5 tok/s127.7 ms187.6 ms200
8143.7 tok/s18 tok/s279.7 ms296.6 ms240
16613.3 tok/s38.3 tok/s410 ms431.4 ms480
Between 8 and 16 streams the aggregate rose 4.27× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q5_K_M) · Q5_K_M — peaks at 517.5 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
175.4 tok/s75.4 tok/s65.1 ms81 ms200
267.5 tok/s33.7 tok/s131.2 ms229.9 ms200
4111.6 tok/s27.9 tok/s293.2 ms399 ms200
8133 tok/s16.6 tok/s632.3 ms819.6 ms240
16517.5 tok/s32.3 tok/s1,032.8 ms1,153.7 ms480
Between 8 and 16 streams the aggregate rose 3.89× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 12B Instruct (Q8_0) · Q8_0 — peaks at 290.2 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
157 tok/s57 tok/s60 ms90.3 ms200
2107.7 tok/s53.8 tok/s102.5 ms150.8 ms200
497.8 tok/s24.4 tok/s315.2 ms379.2 ms200
8290.2 tok/s36.3 tok/s388 ms401.5 ms240
Between 4 and 8 streams the aggregate rose 2.97× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.

Concurrency was capped by VRAM on this card, not by compute.

Mistral Small 3.2 24B Instruct · Q4_K_M — peaks at 437.1 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
153.6 tok/s53.6 tok/s61.2 ms64.1 ms200
250.6 tok/s25.3 tok/s116.2 ms142.8 ms200
481.8 tok/s20.4 tok/s222.8 ms254 ms200
895.1 tok/s11.9 tok/s454.8 ms498.9 ms240
16437.1 tok/s27.3 tok/s764.7 ms883.2 ms480
Between 8 and 16 streams the aggregate rose 4.60× while the stream count rose 2.00×. Batching cannot accelerate — each added stream shares weight reads already happening — so this pair is a measurement artefact, reported rather than smoothed.
Gemma 3 27B Instruct · Q4_K_M — peaks at 110.1 tok/s aggregate
StreamsAggregatePer streamTTFT p50TTFT p95SamplesFailed
141.9 tok/s41.9 tok/s112.9 ms143.3 ms200
272.5 tok/s36.3 tok/s266.2 ms316 ms200
462.9 tok/s15.7 tok/s417.8 ms610 ms200
8110.1 tok/s13.8 tok/s683.8 ms773.5 ms240

Concurrency was capped by VRAM on this card, not by compute.

ONE FLAG

Flash attention, measured at depth

LM Studio and Ollama turn this on by default. Whether it helps, and by how much, depends on the model and only shows up once the KV cache is real — which is why it is measured at depth rather than on an empty cache.

gpt-oss 20B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off191.52,748.111.7 GB3,968
on218.24,199.911.3 GB3,968
Qwen3 30B A3B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off141.81,749.518.2 GB3,968
on186.12,799.218.1 GB3,968
Qwen3 4B Instruct 2507 · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off151.82,756.73.6 GB3,968
on177.65,283.23.5 GB3,968
GLM 4.7 Flash · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off581,499.817.8 GB3,968
on131.11,93017.7 GB3,968
Llama 3.1 8B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off119.32,5485.6 GB3,968
on133.14,101.35.4 GB3,968
Qwen3 8B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off111.32,2435.7 GB3,968
on125.53,6685.5 GB3,968
Qwen3.5 9B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off118.23,211.45.8 GB3,968
on121.63,458.65.7 GB3,968
Gemma 3 12B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off782,3308.7 GB3,968
on81.32,764.88.3 GB3,968
Qwen3 14B · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off68.31,570.29.5 GB3,968
on772,254.89.2 GB3,968
Mistral Small 3.2 24B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off48.81,303.614.4 GB3,968
on51.91,703.514.2 GB3,968
Gemma 3 27B Instruct · Q4_K_M
-faDecode tok/sPrefill tok/sPeak VRAMDepth
off39.61,136.317.3 GB3,968
on40.91,401.717.2 GB3,968
WATTS AND MONEY

What a million tokens costs

Power sampled at the card during generation, not the TDP on the spec sheet. Cost per million tokens uses the rate this pod was billed at; on your own hardware the same energy figures apply against your electricity price.

ModelQuantAvg WPeak WPeak °Ctok/WkWh/Mtok$/Mtok single$/Mtok batched
gpt-oss 20BQ4_K_M248348570.920.302$0.61$0.23
Qwen3 30B A3B Instruct 2507Q4_K_M237349570.880.317$0.67$0.18
Qwen3 4B Instruct 2507Q4_K_M300347590.690.404$0.67$0.13
GLM 4.7 FlashQ4_K_M261349580.570.486$0.93$0.25
Llama 3.1 8B InstructQ4_K_M297348570.490.568$0.96$0.15
Qwen3 8BQ4_K_M305350600.450.613$1.01$0.16
Qwen3 8B (Q5_K_M)Q5_K_M304348620.410.682$1.12$0.17
Qwen3.5 9BQ4_K_M302350610.410.678$1.12$0.27
Qwen3 8B (Q3_K_M)Q3_K_M311348630.380.725$1.17$0.16
Qwen3 8B (Q6_K)Q6_K306349620.350.803$1.31$0.18
Qwen3 8B (Q8_0)Q8_0307350620.30.918$1.49$0.19
Gemma 3 12B InstructQ4_K_M311349620.280.999$1.61$0.25
Qwen3 14BQ4_K_M310348620.261.055$1.7$0.23
Gemma 3 12B Instruct (Q5_K_M)Q5_K_M312348630.251.115$1.79$0.27
Gemma 3 12B Instruct (Q8_0)Q8_0312350630.191.491$2.39$0.48
Mistral Small 3.2 24B InstructQ4_K_M312349630.171.605$2.57$0.32
Gemma 3 27B InstructQ4_K_M316349640.142.059$3.26$1.26
THE MACHINE

What this ran on, in full

A measurement is a claim about a machine, and a machine nobody described is a claim nobody can check. Fields that need root — the DMI table, which a container does not have — say so rather than being dropped: a missing row and an unreadable one mean different things.

PropertyValue
CPUAMD EPYC 7H12 64-Core Processor
Cores128 physical / 256 logical
Cores this process could use27
Threads given to llama-bench27
NUMA nodes2
RAM1008 GB
Kernel6.8.0-107-generic
OSUbuntu 24.04.1 LTS
SystemSupermicro AS -4124GS-TNR-03-WA004
BIOS3.6 03/04/2026
Containerisedyes
Disksmd0 ? 7681 GB, nvme0n1 KINGSTON SKC3000D4096G 4097 GB, nvme1n1 KINGSTON SKC3000D4096G 4097 GB, nvme2n1 KINGSTON SKC3000D4096G 4097 GB, nvme3n1 KINGSTON SKC3000D4096G 4097 GB, sda HFS960G32FEH-7A1 960 GB
GPUVRAMPCIe linkPower limitMax SM / mem clock
0: NVIDIA GeForce RTX 309024.0 GBgen 2 ×16 (of gen 4 ×16)350 W2,100 / 9,751 MHz
▸ WANT ONE FOR YOUR MACHINE?

The NVIDIA GeForce RTX 3090 was rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.

You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.

ASK ABOUT A REPORT →
fitmyllm@gmail.com
HOW THIS WAS PRODUCED

A pod was rented from runpod, llama.cpp b10156 was installed, each model was pulled from Hugging Face, loaded, and timed over 20 runs for latency. The machine terminated itself when the sweep ended. The harness is in bench_lab/report/ and the raw JSON behind this page is served at /api/rig/rtx-3090-24gb — so anything here can be checked against the source rather than taken on trust. How the estimated numbers work →