GTX 760 OEM Rebrand holds 25 of the models in our catalogue and is, in practice, a Q4_K_M card — the largest it takes is Falcon-H1 1.5B at Q4_K_M.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA GeForce GTX 760 OEM Rebrand — 2 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 2 GB GDDR5
- BANDWIDTH
- 179.2 GB/s
- FP16 COMPUTE
- 2 TFLOPS
- FP32 COMPUTE
- 2 TFLOPS
- CUDA CORES
- 1,152
- TDP
- 130 W
- ARCHITECTURE
- Kepler
With 2 GB VRAM and 179.2 GB/s bandwidth, this GPU handles models up to 0.81B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~20 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| GPT-2 Large 774M | 0.81B | 1.0 GB | 177 | 5.6 |
| Qwen3 0.6B | 0.75B | 0.9 GB | 191 | 19.1 |
| LFM2 700M | 0.74B | 0.9 GB | 194 | 50.4 |
| Qwen 1.5 0.5B | 0.62B | 0.9 GB | 231 | 9.9 |
| Qwen3-Embedding 0.6B | 0.6B | 0.9 GB | 239 | — |
| Falcon-H1R Tiny 0.6B | 0.6B | 0.9 GB | 239 | — |
| Falcon Perception 0.6B | 0.6B | 0.9 GB | 239 | — |
| BGE-M3 | 0.568B | 0.8 GB | 252 | 63.0 |
| Snowflake Arctic Embed L v2.0 | 0.568B | 0.8 GB | 252 | — |
| Falcon-H1 0.5B | 0.52B | 0.8 GB | 276 | 41.7 |
| Qwen 2.5 0.5B | 0.5B | 0.8 GB | 287 | 19.4 |
| SmolVLM 500M | 0.5B | 0.8 GB | 287 | — |
| GPT-2 Medium 345M | 0.38B | 0.7 GB | 377 | 5.9 |
| SmolLM2 360M | 0.36B | 0.7 GB | 398 | 8.2 |
| LFM2 350M | 0.35B | 0.7 GB | 410 | 46.3 |
| bge-large-en-v1.5 335M | 0.335B | 0.7 GB | 428 | 62.3 |
| mxbai-embed-large-v1 | 0.335B | 0.7 GB | 428 | 64.7 |
| Snowflake Arctic Embed L | 0.335B | 0.7 GB | 428 | 56.0 |
| Snowflake Arctic Embed M v2.0 | 0.305B | 0.7 GB | 470 | — |
| Gemma 3 270M | 0.27B | 0.7 GB | 531 | 12.6 |