M3 Pro (18GB) holds 248 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is Ling-lite 16.8B at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
Apple M3 Pro (18GB) — 18 GB unified, 12 GB usable.
- BRAND
- Apple
- UNIFIED MEMORY
- 18 GB
- USABLE BY A MODEL
- 12 GB
- BANDWIDTH
- 150 GB/s
- FP16 COMPUTE
- 7.4 TFLOPS
- FP32 COMPUTE
- 7.4 TFLOPS
- TDP
- 30 W
- ARCHITECTURE
- M3 Pro
- MSRP
- $1599
With 12 GB of its 18 GB reaching a model and 150 GB/s bandwidth, this machine handles models up to 14.8B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~17 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| DeepSeek R1 Distill Qwen 14B | 14.8B | 9.5 GB | 9 | 43.9 |
| DeepCoder 14B | 14.8B | 9.5 GB | 9 | 38.7 |
| Qwen2.5-Coder-14B | 14.8B | 9.5 GB | 9 | 41.3 |
| Qwen2.5-14B | 14.8B | 9.5 GB | 9 | 41.3 |
| Qwen3 14B | 14.8B | 9.5 GB | 9 | 45.7 |
| phi-4 14B | 14.66B | 9.4 GB | 9 | 33.7 |
| Phi-4-reasoning 14B | 14.66B | 9.4 GB | 9 | 33.7 |
| Phi-4-reasoning-plus 14B | 14.66B | 9.4 GB | 9 | 75.5 |
| Phi-3-medium-14b | 14B | 9.0 GB | 10 | 33.7 |
| Qwen 1.5 14B | 14B | 9.0 GB | 10 | 41.3 |
| Ministral 3 14B Reasoning | 13.95B | 9.0 GB | 10 | — |
| Baichuan2 13B | 13B | 8.4 GB | 10 | 23.6 |
| Llama 2 13B | 13B | 8.4 GB | 10 | 17.2 |
| CodeLlama 13B | 13B | 8.4 GB | 10 | 19.7 |
| Vicuna 13B | 13B | 8.4 GB | 10 | 11.8 |
| LLaMA 1 13B | 13B | 8.4 GB | 10 | 32.9 |
| OPT 13B | 13B | 8.4 GB | 10 | 35.8 |
| Orca 2 13B | 13B | 8.4 GB | 10 | 25.4 |
| WizardCoder Python 13B | 13B | 8.4 GB | 10 | 60.1 |
| WizardLM 13B | 13B | 8.4 GB | 10 | 19.5 |