Average speeds at Q4 quantization. Actual performance varies by model architecture and context length.
3B
—
1.7GB NEEDED
7B
—
3.9GB NEEDED
14B
—
7.9GB NEEDED
32B
—
18.0GB NEEDED
70B
—
39.4GB NEEDED
▸ MEASURED RIG REPORTS
We rent the machine and time every model on it: decode, VRAM peak, concurrency, watts and cost per million tokens. Yours may already be one of them — and two are free to read in full.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
▸ DEVICE UNDER TEST
NVIDIA GeForce 210 OEM — 1 GB VRAM.
▸ GEFORCE 210 OEM SPEC
BRAND
NVIDIA
VRAM
1 GB DDR2
BANDWIDTH
6.4 GB/s
FP16 COMPUTE
0.1 TFLOPS
FP32 COMPUTE
0.1 TFLOPS
CUDA CORES
16
TDP
31 W
ARCHITECTURE
Tesla 2.0
▸ AI CAPABILITY
4/ 449 models @ Q4
With 1 GB VRAM and 6.4 GB/s bandwidth, this GPU handles models up to 0.14B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~1 tok/s.