Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI).
chatcodingreasoningmultilingualtool_use
80B
Parameters (3B active)
256K
Context length
20
Benchmarks
16
Quantizations
0
Architecture
MoE
Released
2025-09-12
Layers
48
KV Heads
2
Head Dim
256
Family
qwen
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_M
2.93
30.1 GB
29.8 + 0.3 KV
low
Q2_K
3.16
32.4 GB
32.1 + 0.3 KV
low
IQ3_XXS
3.25
33.3 GB
33.0 + 0.3 KV
low
IQ3_XS
3.5
35.8 GB
35.5 + 0.3 KV
low
Q3_K_S
3.64
37.2 GB
36.9 + 0.3 KV
low
IQ3_M
3.76
38.4 GB
38.1 + 0.3 KV
low
Q3_K_M
4
40.8 GB
40.5 + 0.3 KV
low
Q3_K_L
4.3
43.8 GB
43.5 + 0.3 KV
moderate
IQ4_XS
4.46
45.4 GB
45.1 + 0.3 KV
moderate
Q4_K_S
4.67
47.5 GB
47.2 + 0.3 KV
moderate
Q4_K_M
4.89
49.7 GB
49.4 + 0.3 KV
good
Q5_K_S
5.57
56.5 GB
56.2 + 0.3 KV
good
Q5_K_M
5.7
57.8 GB
57.5 + 0.3 KV
good
Q6_K
6.56
66.4 GB
66.1 + 0.3 KV
excellent
Q8_0
8.5
85.8 GB
85.5 + 0.3 KV
lossless
FP16
16
160.8 GB
160.5 + 0.3 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.