FitMyLLM
▸ MEASURED

Rig reports. Nothing estimated.

If you want to buy hardware with your eyes open, the thing worth knowing is what it can actually do. Each report is a machine we rented, with every model loaded and timed on it.

FIND YOUR CARD18 OF 128 CARDS MEASURED
MEASURED · READ OR BUY
  • NVIDIA GeForce RTX 4090
    289 tok/s on gpt-oss 20B · measured 2026-08-01
    FREE →
  • NVIDIA GeForce RTX 3090
    227 tok/s on gpt-oss 20B · measured 2026-07-31
    FREE →
  • 1x L40S 48GB
    231 tok/s on gpt-oss 20B · measured 2026-08-30
    PRO · $25 →
  • 1x RTX A6000 48GB
    206 tok/s on gpt-oss 20B · measured 2026-08-30
    PRO · $25 →
  • 2x A40 48GB
    180 tok/s on gpt-oss 20B · measured 2026-08-30
    PRO · $25 →
  • NVIDIA GeForce RTX 3060 Ti
    123 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
    LIGHT · $9 →
  • NVIDIA GeForce RTX 3070
    126 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
    LIGHT · $9 →
  • NVIDIA GeForce RTX 5060 Ti 16 GB
    150 tok/s on gpt-oss 20B · measured 2026-08-29
    PRO · $25 →
  • NVIDIA GeForce RTX 4070
    151 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
    LIGHT · $9 →
  • NVIDIA GeForce RTX 5070
    186 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
    LIGHT · $9 →
  • NVIDIA GeForce RTX 3060 12 GB
    98 tok/s on Qwen3 4B Instruct 2507 · measured 2026-08-29
    LIGHT · $9 →
  • 1x H100 SXM 80GB
    278 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
    PRO · $25 →
  • 2x L40S
    144 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
    PRO · $25 →
  • 1x H200 141GB
    280 tok/s on Mistral 7B Instruct v0.3 · measured 2026-08-28
    PRO · $25 →
  • NVIDIA RTX 4000 Ada 20GB
    127 tok/s on Qwen3 30B A3B Instruct 2507 · measured 2026-08-25
    PRO · $25 →
  • 1x NVIDIA L4
    92 tok/s on gpt-oss 20B · measured 2026-08-25
    PRO · $25 →
  • 1x NVIDIA A40
    182 tok/s on gpt-oss 20B · measured 2026-08-25
    PRO · $25 →
  • NVIDIA GeForce RTX 5090
    427 tok/s on gpt-oss 20B · measured 2026-08-21
    PRO · $25 →
NOT MEASURED YET · WHAT VISITORS HAVE, MOST COMMON FIRST
▸ WANT ONE FOR YOUR MACHINE?

These machines were rented and measured because we wanted the numbers. If you need the same for hardware we have not run — your models, your context lengths, your concurrency ceiling, your quantisations — that is work we take on commission.

You would get exactly what is on this page: every measurement, the raw JSON, the host it ran on, the build it was measured against — and the results that do not flatter anyone, because the alternative is a report nobody can cite. Tell us the machine and the workload and we will say whether it is worth measuring at all.

ASK ABOUT A REPORT →
fitmyllm@gmail.com
WHAT WE MEASURE, ON EVERY RIG7,349 FIGURES ACROSS 18 MACHINES
DECODE, SINGLE STREAM

Tokens per second generating with the model fully resident on the GPU. Median of several independent runs, each from a cold load.

PREFILL

Tokens per second processing the prompt — the half of latency that a decode figure does not describe.

TTFT P50 AND P95

Time to first token at the median and the 95th percentile. p95 is what an interactive client actually experiences.

PEAK VRAM

The maximum the driver reported during the run, not the weight size. The gap between them is the runtime’s own allocation, and it is why a model fits on paper and not on the card.

QUANTISATION LADDER

The same weights at several rungs on the same card: what each extra bit per weight costs in tokens per second and in VRAM.

DECODE AT DEPTH

Speed with the KV cache already filled to 4K, 8K, 16K, 32K and beyond — not the empty-cache figure, and with the depth at which the model stops loading.

OFFLOAD CURVE

Tokens per second at each fraction of the model kept resident, for the case where it does not fit and layers cross PCIe every token.

CONCURRENCY

Aggregate throughput, per-stream throughput and TTFT p95 at 1, 2, 4, 8 and 16 streams, measured over the seconds the server was actually saturated.

ACCURACY AFTER QUANTISATION

Deterministic probes at temperature 0 with a programmatic check, plus a separate looping test — because a damaged quantisation posts a higher token rate, not a lower one.

POWER AND COST

Watts sampled at the card during generation, and cost per million tokens at the rate the machine was billed, single stream and batched.

FLASH ATTENTION

On and off, at depth. Both LM Studio and Ollama enable it by default; whether it helps depends on the model and the cache size.

FITTED ACROSS THE RUN

The card’s decode constants and implied bandwidth, the fixed VRAM overhead before any weights, and how TTFT scales with concurrency — each with its n and its correlation.

Every speed is the median of several independent processes, each of which reloaded the model from cold; the runs and their spread are printed. Models that failed to load are listed with the stage they failed at, and each report carries the raw JSON it was rendered from.

▸ CLOUD GPU MARKET

September 2026what rent did.

Computed, not written. Every figure comes out of the rates we record from RunPod, Vast.ai and DataCrunch each morning — which means it can be regenerated for any past month and come out the same, and that it cannot exist for a month we did not watch. Free, like the rest of this page except the measurements above.

NOT ENOUGH RECORD YET

September 2026 has 7 days recorded, across 103 cards. A report needs 14.

Below that, one provider outage or a single odd weekend moves every figure in the document and nothing on the page would show it. The record grows by a day each morning, and this becomes a report on its own.

§ ARCHIVE1 EARLIER MONTH
August 2026
WHY THIS EXISTS

Every VRAM calculator on the internet runs the same roofline we do and tells you to go validate it yourself. We went and validated it. The measurements here are also what corrects the estimator: the context curve, the offload curve and the KV-cache formula on this site were all rewritten from these files after they disagreed with the machine. How the estimates work → What a commissioned report contains →