Average speeds at Q4 quantization. Actual performance varies by model architecture and context length.
3B
—
1.7GB NEEDED
7B
—
3.9GB NEEDED
14B
—
7.9GB NEEDED
32B
—
18.0GB NEEDED
70B
—
39.4GB NEEDED
▸ MEASURED RIG REPORTS
We rent the machine and time every model on it: decode, VRAM peak, concurrency, watts and cost per million tokens. Yours may already be one of them — and two are free to read in full.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
▸ DEVICE UNDER TEST
AMD Radeon HD 6390 — 1 GB VRAM.
▸ RADEON HD 6390 SPEC
BRAND
AMD
VRAM
1 GB DDR2
BANDWIDTH
16 GB/s
FP16 COMPUTE
0.4 TFLOPS
FP32 COMPUTE
0.4 TFLOPS
STREAM PROCESSORS
320
TDP
39 W
ARCHITECTURE
TeraScale 2
▸ AI CAPABILITY
4/ 449 models @ Q4
With 1 GB VRAM and 16 GB/s bandwidth, this GPU handles models up to 0.14B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~2 tok/s.