1x H100 SXM 80GB
Before you spend $25,000 on this card, find out if it does what you need.
6 models loaded and timed on this machine, in one session, on one llama.cpp build. Read a whole one first if you like: the RTX 4090 and RTX 3090 reports are open in full.
One payment for this report, not a subscription. It stays yours, and it does not expire. You will be asked to sign in first, so the report stays attached to your account.
- Does Qwen2.5 72B Instruct run on it?
- It did, at Q4_K_M — 45.6 GB of 80 GB at peak
- And how fast, model by model?
- Decode, prefill and TTFT p50/p95 on 6 models
- And when a model is bigger than the card?
- tok/s at each resident fraction, 3 models spilled into system RAM
- How many people can use it at once?
- 1 → 16 streams: aggregate tok/s, per-stream tok/s and TTFT p95
- Does the quantisation cost answers?
- Probed at temperature 0 on 6 models, scored per category
- What does it cost to keep running?
- Watts at the card, and dollars per million tokens, single stream and batched
Decode only, single stream, for the fastest 5. The remaining 1 models, and every other measurement, are listed below.
What is in this report
- · Llama 3.1 8B Instruct · Q4_K_M
- · Mistral 7B Instruct v0.3 · Q4_K_M
- · Phi-3 Medium 14B Instruct · Q4_K_M
- · Qwen2.5 32B Instruct · Q4_K_M
- × Mixtral 8x7B Instruct v0.1 · Q4_K_M
- · Llama 3.3 70B Instruct · Q4_K_M
- · Qwen2.5 72B Instruct · Q4_K_M
Rows marked × failed to load or to run on this card. They are in the report with the stage they failed at, because a model that does not run is a result.
EVERY FIELD, MEASUREMENT BY MEASUREMENT
- · Decode, single stream, tok/s
- · Prefill, tok/s
- · TTFT p50 and p95, ms
- · Peak VRAM reported by the driver
- · Weight size on disk, and load time
- · Rungs measured: Q4_K_M
- · One rung per model in this run
- · 3 models run at several resident fractions
- · tok/s at each fraction, against the fully resident rate
- · Streams: 1, 2, 4, 8, 16
- · Aggregate tok/s, per-stream tok/s and TTFT p95 at each level
- · 6 models load-tested
- · 6 models probed at temperature 0
- · Score per category, and a separate looping check
- · Watts sampled at the card for 6 models
- · Cost per million tokens, single stream and batched
- · The card's decode constants: fixed ms per token, and implied bandwidth
- · Fixed VRAM overhead before the weights — the intercept a fit calculator assumes is zero
- · TTFT p95 against concurrent streams
- · Each with its n and its correlation; a weak fit is quoted without a line
- · GPU: NVIDIA H100 80GB HBM3
- · llama.cpp build b10156
- · Driver 580.126.09
- · Run finished 2026-08-28
- · 0.7 hours on the machine
- · This page, in full, at this URL
- · A PDF with the same tables and charts
- · The underlying JSON, through the API, once the report is yours
- · No expiry, and no subscription to keep reading it
Measured on a rented machine, not estimated and not supplied by a vendor. How the runs are taken, and what each figure does and does not cover, is on the methodology page.