Quadro 5000 holds 55 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is StarCoder2 3B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA Quadro 5000 — 3 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 3 GB GDDR5
- BANDWIDTH
- 120 GB/s
- FP16 COMPUTE
- 0.7 TFLOPS
- FP32 COMPUTE
- 0.7 TFLOPS
- CUDA CORES
- 352
- TDP
- 152 W
- ARCHITECTURE
- Fermi
With 3 GB VRAM and 120 GB/s bandwidth, this GPU handles models up to 2.4B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~14 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| EXAONE Deep 2.4B | 2.4B | 2.0 GB | 40 | 27.1 |
| Qwen3 1.7B | 2.03B | 1.7 GB | 47 | 30.1 |
| OpenCoder 1.5B Instruct | 1.91B | 1.7 GB | 50 | 47.8 |
| InternLM2 1B | 1.89B | 1.6 GB | 51 | — |
| Qwen 1.5 1.8B | 1.8B | 1.6 GB | 53 | 19.6 |
| SmolLM2 1.7B | 1.71B | 1.5 GB | 56 | 16.2 |
| Falcon3-1B | 1.67B | 1.5 GB | 57 | 42.0 |
| GPT-2 XL 1.5B | 1.61B | 1.5 GB | 60 | 5.1 |
| stablelm-2-1_6b | 1.6B | 1.5 GB | 60 | 9.5 |
| LFM2-VL 1.6B | 1.6B | 1.5 GB | 60 | 39.7 |
| Falcon-H1 1.5B | 1.55B | 1.4 GB | 62 | 43.8 |
| Qwen2.5-Coder-1.5B | 1.5B | 1.4 GB | 64 | 19.6 |
| Qwen2 Math 1.5B | 1.5B | 1.4 GB | 64 | 19.6 |
| Qwen 2.5 1.5B | 1.5B | 1.4 GB | 64 | 30.2 |
| Yi Coder 1.5B | 1.5B | 1.4 GB | 64 | 14.6 |
| Stella en 1.5B v5 | 1.5B | 1.4 GB | 64 | — |
| Phi-1 1.3B | 1.42B | 1.4 GB | 68 | 7.2 |
| Phi-1.5 1.3B | 1.42B | 1.4 GB | 68 | 7.2 |
| DeepSeek Coder 1.3B | 1.35B | 1.3 GB | 71 | 16.8 |
| EXAONE-4.0-1.2B | 1.3B | 1.3 GB | 74 | 18.9 |