RTX A2000 Mobile 8 GB holds 201 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is Falcon3-10B at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA RTX A2000 Mobile 8 GB — 8 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 8 GB GDDR6
- BANDWIDTH
- 224 GB/s
- FP16 COMPUTE
- 8.3 TFLOPS
- FP32 COMPUTE
- 8.3 TFLOPS
- CUDA CORES
- 2,560
- TENSOR CORES
- 80
- TDP
- 95 W
- ARCHITECTURE
- Ampere
With 8 GB VRAM and 224 GB/s bandwidth, this GPU handles models up to 9.63B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~26 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| RecurrentGemma 9B | 9.63B | 6.4 GB | 19 | 35.0 |
| glm-4-9b | 9.4B | 6.2 GB | 19 | 20.5 |
| gemma-2-9b | 9.2B | 6.1 GB | 19 | 30.2 |
| Yi 1.5 9B | 9B | 6.0 GB | 20 | 30.3 |
| Yi Coder 9B | 9B | 6.0 GB | 20 | 35.8 |
| Ministral 3 8B | 8.92B | 5.9 GB | 20 | 25.7 |
| Ministral 3 8B Reasoning | 8.92B | 5.9 GB | 20 | — |
| NVIDIA-Nemotron-Nano-9B-v2 | 8.9B | 5.9 GB | 20 | 44.2 |
| InternLM3 8B Instruct | 8.8B | 5.9 GB | 20 | 38.7 |
| Gemma 1 7B | 8.54B | 5.7 GB | 21 | 24.7 |
| CodeGemma 7B | 8.54B | 5.7 GB | 21 | 40.2 |
| LFM2 8B A1B | 8.3B | 5.6 GB | 119 | 24.3 |
| Seed-Coder 8B Instruct | 8.25B | 5.5 GB | 22 | 34.1 |
| Seed-Coder 8B Reasoning | 8.25B | 5.5 GB | 22 | 32.9 |
| DeepSeek R1-0528 Qwen3 8B | 8.2B | 5.5 GB | 22 | 36.3 |
| Qwen3-8B | 8.2B | 5.5 GB | 22 | 43.3 |
| Granite 3.0 8B | 8.17B | 5.5 GB | 22 | 36.4 |
| Granite 3.1 8B | 8.17B | 5.5 GB | 22 | 38.6 |
| Command-R7B | 8.03B | 5.4 GB | 22 | 35.3 |
| Aya Expanse 8B | 8B | 5.4 GB | 22 | 27.8 |