RX 7900 XT holds 276 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is Qwen3.5-27B at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
AMD RX 7900 XT — 20 GB VRAM.
- BRAND
- AMD
- VRAM
- 20 GB GDDR6
- BANDWIDTH
- 800 GB/s
- FP16 COMPUTE
- 103 TFLOPS
- FP32 COMPUTE
- 52 TFLOPS
- STREAM PROCESSORS
- 5,376
- TDP
- 315 W
- ARCHITECTURE
- RDNA3
- MSRP
- $849
With 20 GB VRAM and 800 GB/s bandwidth, this GPU handles models up to 24B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~91 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Mistral-Small-24B | 24B | 15.2 GB | 30 | 25.0 |
| Mistral-Small-3.1-24B | 24B | 15.2 GB | 30 | 28.8 |
| Magistral Small 24B | 24B | 15.2 GB | 30 | 47.0 |
| Devstral Small 2 24B | 24B | 15.2 GB | 30 | 33.4 |
| Mistral-Small-3.2-24B | 24B | 15.2 GB | 30 | 44.4 |
| LFM2 24B A2B | 24B | 15.2 GB | 356 | 19.1 |
| Devstral Small 22B | 23.57B | 14.9 GB | 30 | 35.5 |
| Codestral 22B | 22.2B | 14.1 GB | 32 | 50.1 |
| Mistral Small 22B | 22.2B | 14.1 GB | 32 | 35.2 |
| SOLAR-Pro 22B | 22.1B | 14.0 GB | 32 | 44.2 |
| ERNIE 4.5 21B A3B | 21.95B | 13.9 GB | 237 | — |
| GPT-OSS 20B | 21B | 13.3 GB | 198 | 52.9 |
| Reka Flash 3 | 21B | 13.3 GB | 34 | 36.2 |
| Reka Flash 3.1 | 21B | 13.3 GB | 34 | 33.4 |
| InternLM2 20B | 19.8B | 12.6 GB | 36 | 45.1 |
| InternLM2.5 20B | 19.8B | 12.6 GB | 36 | 50.9 |
| Ling-lite 16.8B | 16.8B | 10.8 GB | 296 | — |
| DeepSeek V2 Lite 16B | 16B | 10.3 GB | 296 | 38.0 |
| StarCoder2 15B | 15.96B | 10.2 GB | 45 | 26.5 |
| DeepSeek-Coder-V2-Lite 15.7B | 15.7B | 10.1 GB | 296 | 43.0 |