RTX 6000D holds 383 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is Mistral Medium 3.5 at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA RTX 6000D — 84 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 84 GB GDDR7
- BANDWIDTH
- 1570 GB/s
- FP16 COMPUTE
- 97 TFLOPS
- FP32 COMPUTE
- 97 TFLOPS
- CUDA CORES
- 19,968
- TENSOR CORES
- 624
- TDP
- 600 W
- ARCHITECTURE
- Blackwell 2.0
- MSRP
- $7500
With 84 GB VRAM and 1570 GB/s bandwidth, this GPU handles models up to 117B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~179 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| GPT-OSS 120B | 117B | 72.0 GB | 246 | 54.1 |
| Command A 111B | 111B | 68.3 GB | 11 | 27.6 |
| GLM 4.5 Air | 110B | 67.7 GB | 105 | 51.0 |
| Qwen 1.5 110B | 110B | 67.7 GB | 11 | 33.4 |
| Llama 4 Scout 17B-16E | 109B | 67.1 GB | 74 | 33.9 |
| Cogito v2 109B MoE | 109B | 67.1 GB | 74 | — |
| Ling 2.6 Flash | 107.49B | 66.2 GB | 170 | 36.8 |
| Sarvam 105B | 105B | 64.7 GB | 12 | 48.0 |
| Command-R+ 104B | 104B | 64.1 GB | 12 | 52.7 |
| Llama-3.2-90B-Vision-Instruct | 90B | 55.5 GB | 14 | 48.5 |
| Hunyuan A13B | 80B | 49.4 GB | 97 | 81.1 |
| Qwen3-Coder-Next | 80B | 49.4 GB | 419 | 43.0 |
| Qwen3-Next 80B A3B | 80B | 49.4 GB | 419 | 49.0 |
| NVLM-D 72B | 79.38B | 49.0 GB | 16 | 48.7 |
| InternVL3 78B | 78B | 48.2 GB | 16 | 80.6 |
| Qwen2.5-72B | 72.7B | 44.9 GB | 17 | 39.7 |
| Qwen2-VL 72B | 72.7B | 44.9 GB | 17 | 55.5 |
| Qwen 1.5 72B | 72B | 44.5 GB | 17 | 49.7 |
| Qwen2 Math 72B | 72B | 44.5 GB | 17 | 49.7 |
| Molmo 72B | 72B | 44.5 GB | 17 | 54.1 |