> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
chatThinkingTool Use
27.8B
Parameters
256K
Context length
21
Benchmarks
10
Quantizations
2.2M
HF downloads
Architecture
Dense
Released
2026-02-24
Layers
64
KV Heads
4
Head Dim
128
Family
qwen
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
Q3_K_M
4
15.9 GB
14.4 + 1.5 KV
low
Q3_K_L
4.3
16.9 GB
15.4 + 1.5 KV
moderate
IQ4_XS
4.46
17.5 GB
16.0 + 1.5 KV
moderate
Q4_K_S
4.67
18.2 GB
16.7 + 1.5 KV
moderate
Q4_K_M
4.89
19.0 GB
17.5 + 1.5 KV
good
Q5_K_S
5.57
21.3 GB
19.8 + 1.5 KV
good
Q5_K_M
5.7
21.8 GB
20.3 + 1.5 KV
good
Q6_K
6.56
24.8 GB
23.3 + 1.5 KV
excellent
Q8_0
8.5
31.5 GB
30.0 + 1.5 KV
lossless
FP16
16
57.6 GB
56.1 + 1.5 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
> This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. > These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
▸ SPEC SHEET
Qwen3.5-27B — 27.8B Dense.
▸ SPECIFICATIONS
PARAMETERS
27.8B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
256K tokens
CAPABILITIES
chat
RELEASE DATE
2026-02-24
PROVIDER
Alibaba
FAMILY
qwen
▸ VRAM REQUIREMENTS
QUANT
BPW
VRAM
QUALITY
Q3_K_M
4
14.4 GB
88%
Q3_K_L
4.3
15.4 GB
90%
IQ4_XS
4.46
16.0 GB
92%
Q4_K_S
4.67
16.7 GB
93%
Q4_K_M
4.89
17.5 GB
94%
Q5_K_S
5.57
19.8 GB
96%
Q5_K_M
5.7
20.3 GB
96%
Q6_K
6.56
23.3 GB
97%
Q8_0
8.5
30.0 GB
100%
FP16
16
56.1 GB
100%
§ 01BENCHMARK SCORES
MMLU-PRO86.1
MATH62.5
IFEval95.0
BBH56.5
MMMU82.3
GPQA11.7
MUSR13.5
BigCodeBench45.0
MMBench92.6
Arena Elo1479.0
GPQA Diamond85.5
HLE24.3
AA Intelligence37.2
AA Coding33.4
LiveCodeBench80.7
SWE-bench72.4
aa_ifbench75.6
aa_terminal_bench32.6
aa_tau293.9
aa_scicode39.5
aa_lcr67.3
§ 02RUN COMMAND
Run Qwen3.5-27B locally with Ollama — needs 17.5 GB VRAM at Q4_K_M: