Tesla M2090 holds 130 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Falcon-H1 7B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA Tesla M2090 — 6 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 6 GB GDDR5
- BANDWIDTH
- 177 GB/s
- FP16 COMPUTE
- 1.3 TFLOPS
- FP32 COMPUTE
- 1.3 TFLOPS
- CUDA CORES
- 512
- TDP
- 250 W
- ARCHITECTURE
- Fermi 2.0
With 6 GB VRAM and 177 GB/s bandwidth, this GPU handles models up to 7B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~20 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Alpaca 7B | 7B | 4.8 GB | 20 | 27.7 |
| Baichuan2 7B | 7B | 4.8 GB | 20 | 21.5 |
| Vicuna 7B | 7B | 4.8 GB | 20 | 22.0 |
| MPT-7B | 7B | 4.8 GB | 20 | 7.8 |
| Orca 2 7B | 7B | 4.8 GB | 20 | 26.1 |
| WizardLM 2 7B | 7B | 4.8 GB | 20 | 26.1 |
| StarCoder2 7B | 7B | 4.8 GB | 20 | 17.0 |
| WizardCoder Python 7B | 7B | 4.8 GB | 20 | 53.7 |
| WizardLM 7B | 7B | 4.8 GB | 20 | 15.5 |
| OLMo 3.1 RLZero 7B Code | 7B | 4.8 GB | 20 | 21.8 |
| OLMo 3.1 RLZero 7B Math | 7B | 4.8 GB | 20 | 21.8 |
| Dolly v2 7B | 6.9B | 4.7 GB | 21 | 7.0 |
| granite-4.0-h-tiny 6.9B | 6.9B | 4.7 GB | 94 | 49.2 |
| Llama 2 7B | 6.74B | 4.6 GB | 21 | 21.1 |
| CodeLlama 7B | 6.74B | 4.6 GB | 21 | 28.1 |
| LLaMA 1 7B | 6.74B | 4.6 GB | 21 | 30.8 |
| DeepSeek Coder 6.7B | 6.7B | 4.6 GB | 21 | 23.6 |
| OPT 6.7B | 6.7B | 4.6 GB | 21 | 18.5 |
| ChatGLM2 6B | 6.24B | 4.3 GB | 23 | 20.7 |
| ChatGLM3 6B | 6.24B | 4.3 GB | 23 | 42.7 |