Average speeds at Q4 quantization. Actual performance varies by model architecture and context length.
3B
—
1.7GB NEEDED
7B
—
3.9GB NEEDED
14B
—
7.9GB NEEDED
32B
—
18.0GB NEEDED
70B
—
39.4GB NEEDED
▸ MEASURED RIG REPORTS
We rent the machine and time every model on it: decode, VRAM peak, concurrency, watts and cost per million tokens. Yours may already be one of them — and two are free to read in full.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
▸ DEVICE UNDER TEST
AMD Radeon HD 5870 — 1 GB VRAM.
▸ RADEON HD 5870 SPEC
BRAND
AMD
VRAM
1 GB GDDR5
BANDWIDTH
153.6 GB/s
FP16 COMPUTE
2.7 TFLOPS
FP32 COMPUTE
2.7 TFLOPS
STREAM PROCESSORS
1,600
TDP
188 W
ARCHITECTURE
TeraScale 2
▸ AI CAPABILITY
4/ 449 models @ Q4
With 1 GB VRAM and 153.6 GB/s bandwidth, this GPU handles models up to 0.14B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~18 tok/s.