Average speeds at Q4 quantization. Actual performance varies by model architecture and context length.
3B
—
1.7GB NEEDED
7B
—
3.9GB NEEDED
14B
—
7.9GB NEEDED
32B
—
18.0GB NEEDED
70B
—
39.4GB NEEDED
▸ MEASURED RIG REPORTS
We rent the machine and time every model on it: decode, VRAM peak, concurrency, watts and cost per million tokens. Yours may already be one of them — and two are free to read in full.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
▸ DEVICE UNDER TEST
AMD Radeon HD 8730 OEM — 1 GB VRAM.
▸ RADEON HD 8730 OEM SPEC
BRAND
AMD
VRAM
1 GB GDDR5
BANDWIDTH
72 GB/s
FP16 COMPUTE
0.6 TFLOPS
FP32 COMPUTE
0.6 TFLOPS
STREAM PROCESSORS
384
TDP
47 W
ARCHITECTURE
GCN 1.0
▸ AI CAPABILITY
4/ 449 models @ Q4
With 1 GB VRAM and 72 GB/s bandwidth, this GPU handles models up to 0.14B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~8 tok/s.