Quadro 6000 SDI holds 130 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Falcon-H1 7B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA Quadro 6000 SDI — 6 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 6 GB GDDR5
- BANDWIDTH
- 143.4 GB/s
- FP16 COMPUTE
- 1 TFLOPS
- FP32 COMPUTE
- 1 TFLOPS
- CUDA CORES
- 448
- TDP
- 231 W
- ARCHITECTURE
- Fermi
With 6 GB VRAM and 143.4 GB/s bandwidth, this GPU handles models up to 7B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~16 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Alpaca 7B | 7B | 4.8 GB | 16 | 27.7 |
| Baichuan2 7B | 7B | 4.8 GB | 16 | 21.5 |
| Vicuna 7B | 7B | 4.8 GB | 16 | 22.0 |
| MPT-7B | 7B | 4.8 GB | 16 | 7.8 |
| Orca 2 7B | 7B | 4.8 GB | 16 | 26.1 |
| WizardLM 2 7B | 7B | 4.8 GB | 16 | 26.1 |
| StarCoder2 7B | 7B | 4.8 GB | 16 | 17.0 |
| WizardCoder Python 7B | 7B | 4.8 GB | 16 | 53.7 |
| WizardLM 7B | 7B | 4.8 GB | 16 | 15.5 |
| OLMo 3.1 RLZero 7B Code | 7B | 4.8 GB | 16 | 21.8 |
| OLMo 3.1 RLZero 7B Math | 7B | 4.8 GB | 16 | 21.8 |
| Dolly v2 7B | 6.9B | 4.7 GB | 17 | 7.0 |
| granite-4.0-h-tiny 6.9B | 6.9B | 4.7 GB | 76 | 49.2 |
| Llama 2 7B | 6.74B | 4.6 GB | 17 | 21.1 |
| CodeLlama 7B | 6.74B | 4.6 GB | 17 | 28.1 |
| LLaMA 1 7B | 6.74B | 4.6 GB | 17 | 30.8 |
| DeepSeek Coder 6.7B | 6.7B | 4.6 GB | 17 | 23.6 |
| OPT 6.7B | 6.7B | 4.6 GB | 17 | 18.5 |
| ChatGLM2 6B | 6.24B | 4.3 GB | 18 | 20.7 |
| ChatGLM3 6B | 6.24B | 4.3 GB | 18 | 42.7 |