M5 (16GB) holds 233 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is Phi-3-medium-14b at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
Apple M5 (16GB) — 16 GB unified, 10.7 GB usable.
- BRAND
- Apple
- UNIFIED MEMORY
- 16 GB
- USABLE BY A MODEL
- 10.7 GB
- BANDWIDTH
- 120 GB/s
- FP16 COMPUTE
- 5.5 TFLOPS
- TDP
- 15 W
- ARCHITECTURE
- M5
With 10.7 GB of its 16 GB reaching a model and 120 GB/s bandwidth, this machine handles models up to 13B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~14 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Baichuan2 13B | 13B | 8.4 GB | 8 | 23.6 |
| Llama 2 13B | 13B | 8.4 GB | 8 | 17.2 |
| CodeLlama 13B | 13B | 8.4 GB | 8 | 19.7 |
| Vicuna 13B | 13B | 8.4 GB | 8 | 11.8 |
| LLaMA 1 13B | 13B | 8.4 GB | 8 | 32.9 |
| OPT 13B | 13B | 8.4 GB | 8 | 35.8 |
| Orca 2 13B | 13B | 8.4 GB | 8 | 25.4 |
| WizardCoder Python 13B | 13B | 8.4 GB | 8 | 60.1 |
| WizardLM 13B | 13B | 8.4 GB | 8 | 19.5 |
| Mistral-Nemo 12.2B | 12.2B | 7.9 GB | 9 | 22.4 |
| Dolly v2 12B | 12B | 7.8 GB | 9 | 6.4 |
| StableLM 2 12B | 12B | 7.8 GB | 9 | 21.3 |
| Falcon2 11B | 11B | 7.2 GB | 10 | 33.2 |
| SOLAR-10.7B | 10.7B | 7.0 GB | 10 | 28.2 |
| Falcon3-10B | 10.3B | 6.8 GB | 10 | 38.2 |
| GLM-4.1V 9B Thinking | 10.29B | 6.8 GB | 10 | — |
| Bamba 9B v2 | 9.78B | 6.5 GB | 11 | 26.1 |
| Qwen 3.5 9B | 9.65B | 6.4 GB | 11 | 50.6 |
| RecurrentGemma 9B | 9.63B | 6.4 GB | 11 | 35.0 |
| glm-4-9b | 9.4B | 6.2 GB | 11 | 20.5 |