L40 CNX holds 293 of the models in our catalogue and is, in practice, a IQ4_XS card — the largest it takes is DeepSeek Coder 33B at IQ4_XS.
Neighbours in memory rather than in price: memory decides whether a card can do the job at all, so two cards of the same size at different prices is the comparison you are making. Ordered by memory, then bandwidth — not by the value column, which is worked out from bandwidth and price alone and therefore rewards a cheap card whatever its software stack does to that bandwidth in practice. Read it as one input, not as a ranking.
NVIDIA L40 CNX — 24 GB VRAM.
- BRAND
- NVIDIA
- VRAM
- 24 GB GDDR6
- BANDWIDTH
- 864 GB/s
- FP16 COMPUTE
- 90 TFLOPS
- FP32 COMPUTE
- 90 TFLOPS
- CUDA CORES
- 18,176
- TENSOR CORES
- 568
- TDP
- 300 W
- ARCHITECTURE
- Ada Lovelace
- MSRP
- $5000
With 24 GB VRAM and 864 GB/s bandwidth, this GPU handles models up to 30.5B parameters.
Speed ≈ bandwidth / model_size × efficiency. A 7B model at Q4 runs at ~99 tok/s.
| MODEL | SIZE | VRAM Q4 | TOK/S | AVG |
|---|---|---|---|---|
| Qwen3 30B A3B | 30.5B | 19.1 GB | 230 | 46.7 |
| Qwen3-Coder 30B-A3B | 30.5B | 19.1 GB | 209 | 36.9 |
| Qwen3-30B-A3B Instruct 2507 | 30.5B | 19.1 GB | 209 | 43.0 |
| MPT-30B | 30B | 18.8 GB | 23 | 26.8 |
| OPT 30B | 30B | 18.8 GB | 23 | 6.3 |
| Qwen3-Omni 30B-A3B | 30B | 18.8 GB | 230 | 40.6 |
| Granite 4.1 30B | 30B | 18.8 GB | 23 | 24.7 |
| TranslateGemma 27B | 28.84B | 18.1 GB | 24 | 38.6 |
| PaliGemma 2 28B | 28B | 17.6 GB | 25 | 38.6 |
| ERNIE 4.5 VL 28B A3B Thinking | 28B | 17.6 GB | 230 | — |
| Qwen3.5-27B | 27.8B | 17.5 GB | 25 | 59.4 |
| Qwen 3.8 27B | 27.78B | 17.5 GB | 25 | 64.6 |
| gemma-3-27b | 27.4B | 17.2 GB | 25 | 27.2 |
| gemma-2-27b | 27.2B | 17.1 GB | 25 | 34.6 |
| Qwen 3.6 27B | 27B | 17.0 GB | 26 | 41.1 |
| Gemma 4 26B A4B | 26B | 16.4 GB | 173 | 47.9 |
| Aria 25B A3.9B | 25.3B | 16.0 GB | 177 | 64.8 |
| Mistral-Small-24B | 24B | 15.2 GB | 29 | 25.0 |
| Mistral-Small-3.1-24B | 24B | 15.2 GB | 29 | 28.8 |
| Magistral Small 24B | 24B | 15.2 GB | 29 | 47.0 |