from mistral_common.tokens.tokenizers.mistral import MistralTokenizer from mistral_common.protocol.instruct.messages import UserMessage from mistral_common.protocol.instruct.request import ChatCompletionRequest
chattool_useTool Use
140.6B
Parameters (39.1B active)
64K
Context length
13
Benchmarks
17
Quantizations
35K
HF downloads
Architecture
MoE
Released
2024-04-16
Layers
56
KV Heads
8
Head Dim
128
Family
mistral
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_XXS
2.38
44.9 GB
42.3 + 2.6 KV
low
IQ2_M
2.93
54.6 GB
52.0 + 2.6 KV
low
Q2_K
3.16
58.7 GB
56.0 + 2.6 KV
low
IQ3_XXS
3.25
60.2 GB
57.6 + 2.6 KV
low
IQ3_XS
3.5
64.6 GB
62.0 + 2.6 KV
low
Q3_K_S
3.64
67.1 GB
64.5 + 2.6 KV
low
IQ3_M
3.76
69.2 GB
66.6 + 2.6 KV
low
Q3_K_M
4
73.4 GB
70.8 + 2.6 KV
low
Q3_K_L
4.3
78.7 GB
76.1 + 2.6 KV
moderate
IQ4_XS
4.46
81.5 GB
78.9 + 2.6 KV
moderate
Q4_K_S
4.67
85.2 GB
82.6 + 2.6 KV
moderate
Q4_K_M
4.89
89.1 GB
86.4 + 2.6 KV
good
Q5_K_S
5.57
101.0 GB
98.4 + 2.6 KV
good
Q5_K_M
5.7
103.3 GB
100.7 + 2.6 KV
good
Q6_K
6.56
118.4 GB
115.8 + 2.6 KV
excellent
Q8_0
8.5
152.5 GB
149.9 + 2.6 KV
lossless
FP16
16
284.3 GB
281.7 + 2.6 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
from mistral_common.tokens.tokenizers.mistral import MistralTokenizer from mistral_common.protocol.instruct.messages import UserMessage from mistral_common.protocol.instruct.request import ChatCompletionRequest
▸ SPEC SHEET
Mixtral-8x22B — 140.6B MoE.
▸ SPECIFICATIONS
PARAMETERS
140.6B (39.1B active)
ARCHITECTURE
Mixture of Experts
CONTEXT LENGTH
64K tokens
CAPABILITIES
chat, tool_use
RELEASE DATE
2024-04-16
PROVIDER
Mistral AI
FAMILY
mistral
▸ VRAM REQUIREMENTS
QUANT
BPW
VRAM
QUALITY
IQ2_XXS
2.38
42.3 GB
65%
IQ2_M
2.93
52.0 GB
75%
Q2_K
3.16
56.0 GB
78%
IQ3_XXS
3.25
57.6 GB
82%
IQ3_XS
3.5
62.0 GB
84%
Q3_K_S
3.64
64.5 GB
85%
IQ3_M
3.76
66.6 GB
86%
Q3_K_M
4
70.8 GB
88%
Q3_K_L
4.3
76.1 GB
90%
IQ4_XS
4.46
78.9 GB
92%
Q4_K_S
4.67
82.6 GB
93%
Q4_K_M
4.89
86.4 GB
94%
Q5_K_S
5.57
98.4 GB
96%
Q5_K_M
5.7
100.7 GB
96%
Q6_K
6.56
115.8 GB
97%
Q8_0
8.5
149.9 GB
100%
FP16
16
281.7 GB
100%
§ 01BENCHMARK SCORES
MMLU-PRO50.7
MATH49.5
IFEval84.0
BBH52.7
GPQA24.9
MUSR17.2
GPQA Diamond33.2
LiveCodeBench14.8
MATH-50054.5
HLE4.1
AA Intelligence9.8
aa_scicode18.8
AIME0.0
§ 02RUN COMMAND
Run Mixtral-8x22B locally with Ollama — needs 86.4 GB VRAM at Q4_K_M: