FitMyLLM

Running LLMs on Apple A14 GPU: What You Can Run

The Apple A14 GPU carries 4GB of unified memory, about 1GB of which reaches a model under macOS, with 25.6 GB/s of memory bandwidth, making it an entry-level option for running local AI models. Here is which models fit, how fast they run, and what quantisation gives the best balance of quality and speed.

Already have it? Find the best model for it. Still choosing? See what a budget buys. Or the full specifications.

Loading the model tables…

Running LLMs on Apple A14 GPU: Complete Guide

The Apple A14 GPU has 4GB of unified memory, of which about 1GB reaches a model under macOS, and 25.6 GB/s memory bandwidth, making it an entry-level option for running local AI models. This guide covers which LLMs fit, expected tok/s performance, and recommended settings.

See the full Apple A14 GPU specs page for detailed specifications, all compatible models, and speed estimates.