> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
chatThinkingTool Use
125.1B
Parameters
256K
Context length
17
Benchmarks
17
Quantizations
621K
HF downloads
Architecture
Dense
Released
2025-06-01
Layers
80
KV Heads
8
Head Dim
128
Family
qwen
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_XXS
2.38
41.5 GB
37.7 + 3.8 KV
low
IQ2_M
2.93
50.1 GB
46.3 + 3.8 KV
low
Q2_K
3.16
53.7 GB
49.9 + 3.8 KV
low
IQ3_XXS
3.25
55.1 GB
51.3 + 3.8 KV
low
IQ3_XS
3.5
59.0 GB
55.2 + 3.8 KV
low
Q3_K_S
3.64
61.2 GB
57.4 + 3.8 KV
low
IQ3_M
3.76
63.0 GB
59.3 + 3.8 KV
low
Q3_K_M
4
66.8 GB
63.0 + 3.8 KV
low
Q3_K_L
4.3
71.5 GB
67.7 + 3.8 KV
moderate
IQ4_XS
4.46
74.0 GB
70.2 + 3.8 KV
moderate
Q4_K_S
4.67
77.3 GB
73.5 + 3.8 KV
moderate
Q4_K_M
4.89
80.7 GB
77.0 + 3.8 KV
good
Q5_K_S
5.57
91.3 GB
87.6 + 3.8 KV
good
Q5_K_M
5.7
93.4 GB
89.6 + 3.8 KV
good
Q6_K
6.56
106.8 GB
103.1 + 3.8 KV
excellent
Q8_0
8.5
137.2 GB
133.4 + 3.8 KV
lossless
FP16
16
254.4 GB
250.7 + 3.8 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
▸ SPEC SHEET
Qwen3.5-122B-A10B — 125.1B Dense.
▸ SPECIFICATIONS
PARAMETERS
125.1B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
256K tokens
CAPABILITIES
chat
RELEASE DATE
2025-06-01
PROVIDER
Alibaba
FAMILY
qwen
▸ VRAM REQUIREMENTS
QUANT
BPW
VRAM
QUALITY
IQ2_XXS
2.38
37.7 GB
65%
IQ2_M
2.93
46.3 GB
75%
Q2_K
3.16
49.9 GB
78%
IQ3_XXS
3.25
51.3 GB
82%
IQ3_XS
3.5
55.2 GB
84%
Q3_K_S
3.64
57.4 GB
85%
IQ3_M
3.76
59.3 GB
86%
Q3_K_M
4
63.0 GB
88%
Q3_K_L
4.3
67.7 GB
90%
IQ4_XS
4.46
70.2 GB
92%
Q4_K_S
4.67
73.5 GB
93%
Q4_K_M
4.89
77.0 GB
94%
Q5_K_S
5.57
87.6 GB
96%
Q5_K_M
5.7
89.6 GB
96%
Q6_K
6.56
103.1 GB
97%
Q8_0
8.5
133.4 GB
100%
FP16
16
250.7 GB
100%
§ 01BENCHMARK SCORES
MMLU-PRO42.5
MATH23.4
IFEval59.4
BBH45.0
GPQA12.2
MUSR16.3
BigCodeBench35.0
Arena Elo1492.0
GPQA Diamond85.7
HLE23.4
AA Intelligence41.6
AA Coding34.7
aa_ifbench75.7
aa_terminal_bench31.1
aa_tau293.6
aa_scicode42.0
aa_lcr66.7
§ 02RUN COMMAND
Run Qwen3.5-122B-A10B locally with Ollama — needs 77.0 GB VRAM at Q4_K_M: