Quadro P3200 Mobile holds 130 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Falcon-H1 7B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA Quadro P3200 Mobile — 6 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 6 GB GDDR5
- BANDWIDTH
- 168.2 GB/s
- FP16 COMPUTE
- 0.1 TFLOPS
- FP32 COMPUTE
- 5.5 TFLOPS
- CUDA CORES
- 1,792
- TDP
- 75 W
- ARCHITECTURE
- Pascal
With 6 GB VRAM and 168.2 GB/s bandwidth, this GPU handles models up to 7B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~19 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Alpaca 7B | 7B | 4.8 GB | 19 | 27.7 |
| Baichuan2 7B | 7B | 4.8 GB | 19 | 21.5 |
| Vicuna 7B | 7B | 4.8 GB | 19 | 22.0 |
| MPT-7B | 7B | 4.8 GB | 19 | 7.8 |
| Orca 2 7B | 7B | 4.8 GB | 19 | 26.1 |
| WizardLM 2 7B | 7B | 4.8 GB | 19 | 26.1 |
| StarCoder2 7B | 7B | 4.8 GB | 19 | 17.0 |
| WizardCoder Python 7B | 7B | 4.8 GB | 19 | 53.7 |
| WizardLM 7B | 7B | 4.8 GB | 19 | 15.5 |
| OLMo 3.1 RLZero 7B Code | 7B | 4.8 GB | 19 | 21.8 |
| OLMo 3.1 RLZero 7B Math | 7B | 4.8 GB | 19 | 21.8 |
| Dolly v2 7B | 6.9B | 4.7 GB | 20 | 7.0 |
| granite-4.0-h-tiny 6.9B | 6.9B | 4.7 GB | 90 | 49.2 |
| Llama 2 7B | 6.74B | 4.6 GB | 20 | 21.1 |
| CodeLlama 7B | 6.74B | 4.6 GB | 20 | 28.1 |
| LLaMA 1 7B | 6.74B | 4.6 GB | 20 | 30.8 |
| DeepSeek Coder 6.7B | 6.7B | 4.6 GB | 20 | 23.6 |
| OPT 6.7B | 6.7B | 4.6 GB | 20 | 18.5 |
| ChatGLM2 6B | 6.24B | 4.3 GB | 22 | 20.7 |
| ChatGLM3 6B | 6.24B | 4.3 GB | 22 | 42.7 |