Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters.
chattool_use
72.7B
Parameters
32K
Context length
15
Benchmarks
16
Quantizations
755K
HF downloads
Architecture
Dense
Released
2024-09-16
Layers
80
KV Heads
8
Head Dim
128
Family
qwen
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_M
2.93
30.9 GB
27.1 + 3.8 KV
low
Q2_K
3.16
33.0 GB
29.2 + 3.8 KV
low
IQ3_XXS
3.25
33.8 GB
30.0 + 3.8 KV
low
IQ3_XS
3.5
36.0 GB
32.3 + 3.8 KV
low
Q3_K_S
3.64
37.3 GB
33.6 + 3.8 KV
low
IQ3_M
3.76
38.4 GB
34.7 + 3.8 KV
low
Q3_K_M
4
40.6 GB
36.8 + 3.8 KV
low
Q3_K_L
4.3
43.3 GB
39.6 + 3.8 KV
moderate
IQ4_XS
4.46
44.8 GB
41.0 + 3.8 KV
moderate
Q4_K_S
4.67
46.7 GB
42.9 + 3.8 KV
moderate
Q4_K_M
4.89
48.7 GB
44.9 + 3.8 KV
good
Q5_K_S
5.57
54.9 GB
51.1 + 3.8 KV
good
Q5_K_M
5.7
56.0 GB
52.3 + 3.8 KV
good
Q6_K
6.56
63.9 GB
60.1 + 3.8 KV
excellent
Q8_0
8.5
81.5 GB
77.7 + 3.8 KV
lossless
FP16
16
149.6 GB
145.9 + 3.8 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters.
▸ SPEC SHEET
Qwen2.5-72B — 72.7B Dense.
▸ SPECIFICATIONS
PARAMETERS
72.7B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
32K tokens
CAPABILITIES
chat, tool_use
RELEASE DATE
2024-09-16
PROVIDER
Alibaba
FAMILY
qwen
▸ VRAM REQUIREMENTS
QUANT
BPW
VRAM
QUALITY
IQ2_M
2.93
27.1 GB
75%
Q2_K
3.16
29.2 GB
78%
IQ3_XXS
3.25
30.0 GB
82%
IQ3_XS
3.5
32.3 GB
84%
Q3_K_S
3.64
33.6 GB
85%
IQ3_M
3.76
34.7 GB
86%
Q3_K_M
4
36.8 GB
88%
Q3_K_L
4.3
39.6 GB
90%
IQ4_XS
4.46
41.0 GB
92%
Q4_K_S
4.67
42.9 GB
93%
Q4_K_M
4.89
44.9 GB
94%
Q5_K_S
5.57
51.1 GB
96%
Q5_K_M
5.7
52.3 GB
96%
Q6_K
6.56
60.1 GB
97%
Q8_0
8.5
77.7 GB
100%
FP16
16
145.9 GB
100%
§ 01BENCHMARK SCORES
HumanEval59.1
MMLU-PRO51.4
MATH59.8
IFEval86.4
BBH61.9
GPQA16.7
MUSR11.7
MBPP61.6
BigCodeBench38.5
Arena Elo1481.0
LiveCodeBench27.6
AIME14.0
MATH-50014.0
GPQA Diamond49.1
HLE4.2
§ 02RUN COMMAND
Run Qwen2.5-72B locally with Ollama — needs 44.9 GB VRAM at Q4_K_M: