A15 GPU 5-core holds 81 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Falcon3-3B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
Apple A15 GPU 5-core — 6 GB unified, 3 GB usable.
- BRAND
- Apple
- UNIFIED MEMORY
- 6 GB
- USABLE BY A MODEL
- 3 GB
- BANDWIDTH
- 34.1 GB/s
- FP16 COMPUTE
- 1.5 TFLOPS
- TDP
- 6 W
- ARCHITECTURE
- Apple GPU 5-core
With 3 GB of its 6 GB reaching a model and 34.1 GB/s bandwidth, this machine handles models up to 3.1B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~4 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Qwen 2.5 3B | 3.1B | 2.4 GB | 10 | 37.2 |
| SmolLM3-3B | 3.1B | 2.4 GB | 10 | 30.5 |
| Ministral 3B | 3B | 2.3 GB | 10 | 29.6 |
| StarCoder2 3B | 3B | 2.3 GB | 10 | 9.5 |
| Granite 4.1 3B | 3B | 2.3 GB | 10 | 16.6 |
| xLAM-2 3B Function-Calling | 3B | 2.3 GB | 10 | — |
| Jamba 2 3B | 3B | 2.3 GB | 10 | — |
| Dolly v2 3B | 2.8B | 2.2 GB | 11 | 5.6 |
| StableLM Zephyr 3B | 2.79B | 2.2 GB | 11 | 14.9 |
| Zephyr 3B | 2.79B | 2.2 GB | 11 | 14.4 |
| OPT 2.7B | 2.7B | 2.1 GB | 11 | 28.0 |
| Phi-2 2.7B | 2.7B | 2.1 GB | 11 | 24.1 |
| Zamba2 2.7B | 2.7B | 2.1 GB | 11 | 48.0 |
| Granite 3.0 2B | 2.63B | 2.1 GB | 12 | 35.8 |
| gemma-2-2b | 2.6B | 2.1 GB | 12 | 22.9 |
| LFM2 2.6B | 2.6B | 2.1 GB | 12 | 16.3 |
| Granite 3.1 2B | 2.53B | 2.0 GB | 12 | 37.8 |
| Granite 3.3 2B | 2.53B | 2.0 GB | 12 | 20.5 |
| Gemma 1 2B | 2.51B | 2.0 GB | 12 | 20.2 |
| CodeGemma 2B | 2.51B | 2.0 GB | 12 | 22.9 |