Qualcomm Adreno 740 holds 262 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is ERNIE 4.5 21B A3B at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
Qualcomm Adreno 740 — 16 GB VRAM.
- BRAND
- Apple
- VRAM
- 16 GB Shared
- BANDWIDTH
- 68.3 GB/s
- FP16 COMPUTE
- 2.2 TFLOPS
- TDP
- 10 W
- ARCHITECTURE
- Adreno 700
With 16 GB VRAM and 68.3 GB/s bandwidth, this GPU handles models up to 19.8B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~8 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| InternLM2 20B | 19.8B | 12.6 GB | 3 | 45.1 |
| InternLM2.5 20B | 19.8B | 12.6 GB | 3 | 50.9 |
| Ling-lite 16.8B | 16.8B | 10.8 GB | 25 | — |
| DeepSeek V2 Lite 16B | 16B | 10.3 GB | 25 | 38.0 |
| StarCoder2 15B | 15.96B | 10.2 GB | 4 | 26.5 |
| DeepSeek-Coder-V2-Lite 15.7B | 15.7B | 10.1 GB | 25 | 43.0 |
| DeepSeek-VL2 Small 16B | 15.7B | 10.1 GB | 25 | 43.1 |
| StarCoder 15B | 15.5B | 10.0 GB | 4 | 21.0 |
| InternVL3 14B | 15.12B | 9.7 GB | 4 | 38.1 |
| Phi-4-reasoning-vision 15B | 15B | 9.7 GB | 4 | 42.8 |
| DeepSeek R1 Distill Qwen 14B | 14.8B | 9.5 GB | 4 | 43.9 |
| DeepCoder 14B | 14.8B | 9.5 GB | 4 | 38.7 |
| Qwen2.5-Coder-14B | 14.8B | 9.5 GB | 4 | 41.3 |
| Qwen2.5-14B | 14.8B | 9.5 GB | 4 | 41.3 |
| Qwen3 14B | 14.8B | 9.5 GB | 4 | 45.7 |
| phi-4 14B | 14.66B | 9.4 GB | 4 | 33.7 |
| Phi-4-reasoning 14B | 14.66B | 9.4 GB | 4 | 33.7 |
| Phi-4-reasoning-plus 14B | 14.66B | 9.4 GB | 4 | 75.5 |
| Ministral 3 14B | 14B | 9.0 GB | 4 | 25.9 |
| Phi-3-medium-14b | 14B | 9.0 GB | 4 | 33.7 |