Average speeds at Q4 quantization. Actual performance varies by model architecture and context length.
3B
—
1.7GB NEEDED
7B
—
3.9GB NEEDED
14B
—
7.9GB NEEDED
32B
—
18.0GB NEEDED
70B
—
39.4GB NEEDED
▸ MEASURED RIG REPORTS
We rent the machine and time every model on it: decode, VRAM peak, concurrency, watts and cost per million tokens. Yours may already be one of them — and two are free to read in full.
Radeon HD 6770M Mac Edition holds 4 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Qwen 3.5 0.8B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
▸ DEVICE UNDER TEST
AMD Radeon HD 6770M Mac Edition — 1 GB VRAM.
▸ RADEON HD 6770M MAC EDITION SPEC
BRAND
AMD
VRAM
1 GB GDDR5
BANDWIDTH
50.8 GB/s
FP16 COMPUTE
0.6 TFLOPS
FP32 COMPUTE
0.6 TFLOPS
STREAM PROCESSORS
480
TDP
35 W
ARCHITECTURE
TeraScale 2
▸ AI CAPABILITY
4/ 449 models @ Q4
With 1 GB VRAM and 50.8 GB/s bandwidth, this GPU handles models up to 0.14B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~6 tok/s.