Molmo is a family of open vision-language models developed by the Allen Institute for AI. Molmo models are trained on PixMo, a dataset of 1 million, highly-curated image-text pairs.
chatvision
72B
Parameters
4K
Context length
1
Benchmarks
16
Quantizations
0
Architecture
Dense
Released
2024-09-24
Layers
80
KV Heads
8
Head Dim
128
Family
other
Quantization Options
Quant
Bits
VRAM @ 4K
Quality
IQ2_M
2.93
26.9 GB
low
Q2_K
3.16
28.9 GB
low
IQ3_XXS
3.25
29.7 GB
low
IQ3_XS
3.5
32.0 GB
low
Q3_K_S
3.64
33.2 GB
low
IQ3_M
3.76
34.3 GB
low
Q3_K_M
4
36.5 GB
low
Q3_K_L
4.3
39.2 GB
moderate
IQ4_XS
4.46
40.6 GB
moderate
Q4_K_S
4.67
42.5 GB
moderate
Q4_K_M
4.89
44.5 GB
good
Q5_K_S
5.57
50.6 GB
good
Q5_K_M
5.7
51.8 GB
good
Q6_K
6.56
59.5 GB
excellent
Q8_0
8.5
77.0 GB
lossless
FP16
16
144.5 GB
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.