RTX 3080 10GB holds 223 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Mistral-Nemo 12.2B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA RTX 3080 10GB — 10 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 10 GB GDDR6X
- BANDWIDTH
- 760 GB/s
- FP16 COMPUTE
- 60 TFLOPS
- FP32 COMPUTE
- 29.8 TFLOPS
- CUDA CORES
- 8,704
- TENSOR CORES
- 272
- TDP
- 320 W
- ARCHITECTURE
- Ampere
- MSRP
- $429
With 10 GB VRAM and 760 GB/s bandwidth, this GPU handles models up to 12.2B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~87 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Mistral-Nemo 12.2B | 12.2B | 7.9 GB | 50 | 22.4 |
| Dolly v2 12B | 12B | 7.8 GB | 51 | 6.4 |
| StableLM 2 12B | 12B | 7.8 GB | 51 | 21.3 |
| Falcon2 11B | 11B | 7.2 GB | 55 | 33.2 |
| SOLAR-10.7B | 10.7B | 7.0 GB | 57 | 28.2 |
| Falcon3-10B | 10.3B | 6.8 GB | 59 | 38.2 |
| Bamba 9B v2 | 9.78B | 6.5 GB | 62 | 26.1 |
| Qwen 3.5 9B | 9.65B | 6.4 GB | 63 | 50.6 |
| RecurrentGemma 9B | 9.63B | 6.4 GB | 63 | 35.0 |
| glm-4-9b | 9.4B | 6.2 GB | 65 | 20.5 |
| MiniCPM-o 4.5 | 9.37B | 6.2 GB | 65 | — |
| gemma-2-9b | 9.2B | 6.1 GB | 66 | 30.2 |
| Yi 1.5 9B | 9B | 6.0 GB | 68 | 30.3 |
| Yi Coder 9B | 9B | 6.0 GB | 68 | 35.8 |
| Ministral 3 8B | 8.92B | 5.9 GB | 68 | 25.7 |
| Ministral 3 8B Reasoning | 8.92B | 5.9 GB | 68 | — |
| NVIDIA-Nemotron-Nano-9B-v2 | 8.9B | 5.9 GB | 68 | 44.2 |
| InternLM3 8B Instruct | 8.8B | 5.9 GB | 69 | 38.7 |
| Qwen3-VL 8B Instruct | 8.77B | 5.8 GB | 69 | 26.4 |
| MiniCPM-V 4.5 | 8.7B | 5.8 GB | 70 | 26.1 |