This family of models performs vision-language and text-only tasks including optical character recognition, multimodal reasoning, localization, common sense reasoning, world knowledge utilization, and coding.
chatvisionreasoning
79.38B
Parameters
32K
Context length
9
Benchmarks
16
Quantizations
0
Architecture
Dense
Released
2024-10-08
Layers
80
KV Heads
8
Head Dim
128
Family
nemotron
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_M
2.93
33.3 GB
29.6 + 3.8 KV
low
Q2_K
3.16
35.6 GB
31.8 + 3.8 KV
low
IQ3_XXS
3.25
36.5 GB
32.7 + 3.8 KV
low
IQ3_XS
3.5
39.0 GB
35.2 + 3.8 KV
low
Q3_K_S
3.64
40.4 GB
36.6 + 3.8 KV
low
IQ3_M
3.76
41.5 GB
37.8 + 3.8 KV
low
Q3_K_M
4
43.9 GB
40.2 + 3.8 KV
low
Q3_K_L
4.3
46.9 GB
43.2 + 3.8 KV
moderate
IQ4_XS
4.46
48.5 GB
44.7 + 3.8 KV
moderate
Q4_K_S
4.67
50.6 GB
46.8 + 3.8 KV
moderate
Q4_K_M
4.89
52.8 GB
49.0 + 3.8 KV
good
Q5_K_S
5.57
59.5 GB
55.8 + 3.8 KV
good
Q5_K_M
5.7
60.8 GB
57.0 + 3.8 KV
good
Q6_K
6.56
69.3 GB
65.6 + 3.8 KV
excellent
Q8_0
8.5
88.6 GB
84.8 + 3.8 KV
lossless
FP16
16
163.0 GB
159.2 + 3.8 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.