FitMyLLM
Alibaba/Mixture of Experts

AlibabaQwen3.8 2.4T A95B

Qwen3.8 2.4T-A95B — the open-weight release behind Qwen3.8-Max, and the first Max-class Qwen published as weights. 2.4T total, 95B active (10 routed + 1 shared of 512 experts), text-only, 256K context. The hosted Max adds vision input, non-thinking mode and 1M context; these weights do not. Custom Qwen license, not Apache.

chatcodingreasoningmultilingualmathagentictool_use
2446.18B
Parameters (95B active)
256K
Context length
6
Benchmarks
17
Quantizations
6K
HF downloads
Architecture
MoE
Released
2026-08-12
Layers
92
KV Heads
4
Head Dim
256
Family
qwen

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ2_XXS2.38
729.3 GB
728.2 + 1.1 KV
low
IQ2_M2.93
897.5 GB
896.4 + 1.1 KV
low
Q2_K3.16
967.8 GB
966.7 + 1.1 KV
low
IQ3_XXS3.25
995.3 GB
994.2 + 1.1 KV
low
IQ3_XS3.5
1071.8 GB
1070.7 + 1.1 KV
low
Q3_K_S3.64
1114.6 GB
1113.5 + 1.1 KV
low
IQ3_M3.76
1151.3 GB
1150.2 + 1.1 KV
low
Q3_K_M4
1224.7 GB
1223.6 + 1.1 KV
low
Q3_K_L4.3
1316.4 GB
1315.3 + 1.1 KV
moderate
IQ4_XS4.46
1365.3 GB
1364.2 + 1.1 KV
moderate
Q4_K_S4.67
1429.5 GB
1428.4 + 1.1 KV
moderate
Q4_K_M4.89
1496.8 GB
1495.7 + 1.1 KV
good
Q5_K_S5.57
1704.7 GB
1703.6 + 1.1 KV
good
Q5_K_M5.7
1744.5 GB
1743.4 + 1.1 KV
good
Q6_K6.56
2007.4 GB
2006.4 + 1.1 KV
excellent
Q8_08.5
2600.6 GB
2599.6 + 1.1 KV
lossless
FP1616
4893.9 GB
4892.8 + 1.1 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Too big for a single GPU — plan a multi-GPU deployment
Even the lightest quant needs ~729 GB. Size GPUs, replicas, TCO and scaling for a production setup. Open in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN QWEN3.8 2.4T A95B NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (6)

GPQA Diamond93.5
AA Long Context75.3
AA Coding71.9
AA Intelligence57.7
SciCode51.6
HLE42.4

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run qwen:2446b-q4_K_M

Tag may need adjustment — check ollama.com/library/qwen for available tags.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

Find the best GPU for Qwen3.8 2.4T A95B

Build Hardware for Qwen3.8 2.4T A95B
▸ SPEC SHEET

Qwen3.8 2.4T A95B2446.18B MoE.

▸ SPECIFICATIONS
PARAMETERS
2446.18B (95B active)
ARCHITECTURE
Mixture of Experts
CONTEXT LENGTH
256K tokens
CAPABILITIES
chat, coding, reasoning, multilingual, math, agentic, tool_use
RELEASE DATE
2026-08-12
PROVIDER
Alibaba
FAMILY
qwen
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ2_XXS2.38728.2 GB65%
IQ2_M2.93896.4 GB75%
Q2_K3.16966.7 GB78%
IQ3_XXS3.25994.2 GB82%
IQ3_XS3.51070.7 GB84%
Q3_K_S3.641113.5 GB85%
IQ3_M3.761150.2 GB86%
Q3_K_M41223.6 GB88%
Q3_K_L4.31315.3 GB90%
IQ4_XS4.461364.2 GB92%
Q4_K_S4.671428.4 GB93%
Q4_K_M4.891495.7 GB94%
Q5_K_S5.571703.6 GB96%
Q5_K_M5.71743.4 GB96%
Q6_K6.562006.4 GB97%
Q8_08.52599.6 GB100%
FP16164892.8 GB100%
§ 01BENCHMARK SCORES
GPQA Diamond93.5
HLE42.4
AA Intelligence57.7
AA Coding71.9
aa_scicode51.6
aa_lcr75.3