Jamba2 Mini is an open source small language model built for enterprise reliability. With 12B active parameters (52B total), it delivers precise question answering without the computational overhead of reasoning models.
chatreasoningtool_usemultilingual
52B
Parameters (12B active)
256K
Context length
6
Benchmarks
14
Quantizations
0
Architecture
MoE
Released
2026-02-15
Layers
32
KV Heads
8
Head Dim
128
Family
jamba
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ3_XXS
3.25
23.1 GB
21.6 + 1.5 KV
low
IQ3_XS
3.5
24.7 GB
23.2 + 1.5 KV
low
Q3_K_S
3.64
25.6 GB
24.1 + 1.5 KV
low
IQ3_M
3.76
26.4 GB
24.9 + 1.5 KV
low
Q3_K_M
4
28.0 GB
26.5 + 1.5 KV
low
Q3_K_L
4.3
29.9 GB
28.4 + 1.5 KV
moderate
IQ4_XS
4.46
31.0 GB
29.5 + 1.5 KV
moderate
Q4_K_S
4.67
32.3 GB
30.8 + 1.5 KV
moderate
Q4_K_M
4.89
33.8 GB
32.3 + 1.5 KV
good
Q5_K_S
5.57
38.2 GB
36.7 + 1.5 KV
good
Q5_K_M
5.7
39.0 GB
37.5 + 1.5 KV
good
Q6_K
6.56
44.6 GB
43.1 + 1.5 KV
excellent
Q8_0
8.5
57.2 GB
55.7 + 1.5 KV
lossless
FP16
16
106.0 GB
104.5 + 1.5 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.