FitMyLLM
▸ MEASURED RIG REPORT$9LIGHT · 6 MODELS

1x H100 SXM 80GB

6 models measured80 GB VRAMllama.cpp b101562026-08-280.7 h on the machine

Before you spend $25,000 on this card, find out if it does what you need.

6 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.

One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.

§ WHAT THIS REPORT SETTLES
Does Qwen2.5 72B Instruct run on it?
It did, at Q4_K_M — 45.6 GB of 80 GB at peak
And how fast, model by model?
Decode, prefill and TTFT p50/p95 on 6 models
And when a model is bigger than the card?
tok/s at each resident fraction, 3 models spilled into system RAM
How many people can use it at once?
1 → 16 streams: aggregate tok/s, per-stream tok/s and TTFT p95
Does the quantisation cost answers?
Probed at temperature 0 on 6 models, scored per category
What does it cost to keep running?
Watts at the card, and dollars per million tokens, single stream and batched
§ PREVIEW · THE FASTEST 5MEASURED, NOT ESTIMATED
Mistral 7B Instruct v0.3 · Q4_K_M278
Llama 3.1 8B Instruct · Q4_K_M263
Phi-3 Medium 14B Instruct · Q4_K_M171
Qwen2.5 32B Instruct · Q4_K_M73.3
Llama 3.3 70B Instruct · Q4_K_M41.0
DECODE TOK/S, SINGLE STREAM

Decode only, single stream, for the fastest 5. The remaining 1 models, and every other measurement, are listed below.

What is in this report

MODELS RUN ON THIS MACHINE6 MEASURED · 1 FAILED
  • · Llama 3.1 8B Instruct · Q4_K_M
  • · Mistral 7B Instruct v0.3 · Q4_K_M
  • · Phi-3 Medium 14B Instruct · Q4_K_M
  • · Qwen2.5 32B Instruct · Q4_K_M
  • × Mixtral 8x7B Instruct v0.1 · Q4_K_M
  • · Llama 3.3 70B Instruct · Q4_K_M
  • · Qwen2.5 72B Instruct · Q4_K_M

Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.

EVERY FIELD, MEASUREMENT BY MEASUREMENT
FOR EVERY MODEL
  • · Decode, single stream, tok/s
  • · Prefill, tok/s
  • · TTFT p50 and p95, ms
  • · Peak VRAM reported by the driver
  • · Weight size on disk, and load time
QUANTISATION
  • · Rungs measured: Q4_K_M
  • · One rung per model in this run
OFFLOAD TO SYSTEM RAM
  • · 3 models run at several resident fractions
  • · tok/s at each fraction, against the fully resident rate
CONCURRENCY
  • · Streams: 1, 2, 4, 8, 16
  • · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
  • · 6 models load-tested
ACCURACY AFTER QUANTISATION
  • · 6 models probed at temperature 0
  • · Score per category, and a separate looping check
POWER AND COST
  • · Watts sampled at the card for 6 models
  • · Cost per million tokens, single stream and batched
FITTED ACROSS THE RUN
  • · The card's decode constants: fixed ms per token, and implied bandwidth
  • · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
  • · TTFT p95 against concurrent streams
  • · Each with its n and its correlation; a weak fit is quoted without a line
PROVENANCE
  • · GPU: NVIDIA H100 80GB HBM3
  • · llama.cpp build b10156
  • · Driver 580.126.09
  • · Run finished 2026-08-28
  • · 0.7 hours on the machine
WHAT YOU GET
  • · This page, in full, at this URL
  • · A PDF with the same tables and charts
  • · The underlying JSON, through the API, once the report is yours
  • · No expiry, and no subscription to keep reading it

Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.