FitMyLLM

Running LLMs on Google Tensor G4 GPU: What You Can Run

The Google Tensor G4 GPU packs 12GB of Shared with 51.2 GB/s of memory bandwidth, making it a capable option for running local AI models. Here is which models fit, how fast they run, and what quantisation gives the best balance of quality and speed.

Already have it? Find the best model for it. Still choosing? See what a budget buys. Or the full specifications.

Loading the model tables…

Running LLMs on Google Tensor G4 GPU: Complete Guide

The Google Tensor G4 GPU has 12GB VRAM and 51.2 GB/s memory bandwidth, making it a capable option for running local AI models. This guide covers which LLMs fit, expected tok/s performance, and recommended settings.

See the full Google Tensor G4 GPU specs page for detailed specifications, all compatible models, and speed estimates.