H100 SXM5 64 GB holds 373 of the models in our catalogue and is, in practice, a Q4_K_S card — the largest it takes is Llama-3.2-90B-Vision-Instruct at Q4_K_S.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA H100 SXM5 64 GB — 64 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 64 GB HBM3
- BANDWIDTH
- 2020 GB/s
- FP16 COMPUTE
- 267.6 TFLOPS
- FP32 COMPUTE
- 66.9 TFLOPS
- CUDA CORES
- 16,896
- TENSOR CORES
- 528
- TDP
- 700 W
- ARCHITECTURE
- Hopper
- MSRP
- $25000
With 64 GB VRAM and 2020 GB/s bandwidth, this GPU handles models up to 80B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~231 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Hunyuan A13B | 80B | 49.4 GB | 124 | 81.1 |
| Qwen3-Coder-Next | 80B | 49.4 GB | 539 | 43.0 |
| Qwen3-Next 80B A3B | 80B | 49.4 GB | 539 | 49.0 |
| NVLM-D 72B | 79.38B | 49.0 GB | 20 | 48.7 |
| InternVL3 78B | 78B | 48.2 GB | 21 | 80.6 |
| Qwen2.5-72B | 72.7B | 44.9 GB | 22 | 39.7 |
| Qwen2-VL 72B | 72.7B | 44.9 GB | 22 | 55.5 |
| Qwen 1.5 72B | 72B | 44.5 GB | 22 | 49.7 |
| Qwen2 Math 72B | 72B | 44.5 GB | 22 | 49.7 |
| Molmo 72B | 72B | 44.5 GB | 22 | 54.1 |
| DeepSeek R1 Distill Llama 70B | 70.6B | 43.6 GB | 23 | 42.4 |
| Llama 3.3 70B | 70.6B | 43.6 GB | 23 | 44.8 |
| Llama 3.1 70B | 70.6B | 43.6 GB | 23 | 33.2 |
| Llama 3 70B | 70.6B | 43.6 GB | 23 | 44.1 |
| Llama-3.1-Nemotron-70B | 70.6B | 43.6 GB | 23 | 43.7 |
| Cogito 70B | 70B | 43.3 GB | 23 | — |
| Llama 2 70B | 70B | 43.3 GB | 23 | 33.4 |
| CodeLlama 70B | 70B | 43.3 GB | 23 | 45.7 |
| Dolphin Llama 3 70B | 70B | 43.3 GB | 23 | 45.7 |
| Tulu 3 70B | 70B | 43.3 GB | 23 | 59.4 |