M3 (8GB) holds 104 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is ChatGLM2 6B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
Apple M3 (8GB) — 8 GB unified, 5 GB usable.
- BRAND
- Apple
- UNIFIED MEMORY
- 8 GB
- USABLE BY A MODEL
- 5 GB
- BANDWIDTH
- 100 GB/s
- FP16 COMPUTE
- 4.1 TFLOPS
- FP32 COMPUTE
- 4.1 TFLOPS
- TDP
- 22 W
- ARCHITECTURE
- M3
- MSRP
- $599
With 5 GB of its 8 GB reaching a model and 100 GB/s bandwidth, this machine handles models up to 4.5B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~11 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| InternLM2 5B | 4.5B | 3.2 GB | 20 | 47.6 |
| Qwen3-VL 4B Instruct | 4.44B | 3.2 GB | 20 | 26.6 |
| MedGemma 1.5 4B | 4.3B | 3.1 GB | 21 | 5.5 |
| Gemma 3 4B | 4.3B | 3.1 GB | 21 | 22.8 |
| Ministral 3 3B Reasoning | 4.25B | 3.1 GB | 21 | — |
| Qwen 1.5 4B | 4B | 2.9 GB | 22 | 12.6 |
| Qwen3 4B | 4B | 2.9 GB | 22 | 40.7 |
| Qwen3-4B Instruct 2507 | 4B | 2.9 GB | 22 | 37.2 |
| Qwen3-Embedding 4B | 4B | 2.9 GB | 22 | — |
| Granite Vision 4.1 4B | 4B | 2.9 GB | 22 | — |
| Nemotron 3 Nano 4B | 3.97B | 2.9 GB | 22 | 32.0 |
| Ministral 3 3B | 3.85B | 2.8 GB | 23 | 21.4 |
| Phi-3.5 Mini 3.8B | 3.82B | 2.8 GB | 23 | 46.6 |
| phi-3-mini-4k 3.8B | 3.8B | 2.8 GB | 23 | 30.5 |
| Phi-4-mini 3.8B | 3.8B | 2.8 GB | 23 | 49.0 |
| Qwen2.5-VL-3B | 3.8B | 2.8 GB | 23 | 29.9 |
| Jina Embeddings v4 | 3.8B | 2.8 GB | 23 | — |
| Cogito 3B | 3.61B | 2.7 GB | 25 | 22.1 |
| Falcon3-3B | 3.23B | 2.5 GB | 28 | 25.7 |
| granite-4.0-h-micro 3.2B | 3.2B | 2.4 GB | 28 | 18.4 |