The Mistral-Nemo-Instruct-2407 Large Language Model (LLM) is an instruct fine-tuned version of the Mistral-Nemo-Base-2407. Trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.
chattool_useTool Use
12.2B
Parameters
128K
Context length
7
Benchmarks
10
Quantizations
0
Architecture
Dense
Released
2024-07-18
Layers
40
KV Heads
8
Head Dim
128
Family
mistral
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
Q3_K_M
4
8.5 GB
6.6 + 1.9 KV
low
Q3_K_L
4.3
8.9 GB
7.0 + 1.9 KV
moderate
IQ4_XS
4.46
9.2 GB
7.3 + 1.9 KV
moderate
Q4_K_S
4.67
9.5 GB
7.6 + 1.9 KV
moderate
Q4_K_M
4.89
9.8 GB
7.9 + 1.9 KV
good
Q5_K_S
5.57
10.9 GB
9.0 + 1.9 KV
good
Q5_K_M
5.7
11.1 GB
9.2 + 1.9 KV
good
Q6_K
6.56
12.4 GB
10.5 + 1.9 KV
excellent
Q8_0
8.5
15.3 GB
13.5 + 1.9 KV
lossless
FP16
16
26.8 GB
24.9 + 1.9 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
The Mistral-Nemo-Instruct-2407 Large Language Model (LLM) is an instruct fine-tuned version of the Mistral-Nemo-Base-2407. Trained jointly by Mistral AI and NVIDIA, it significantly outperforms existing models smaller or similar in size.
▸ SPEC SHEET
Mistral-Nemo 12.2B — 12.2B Dense.
▸ SPECIFICATIONS
PARAMETERS
12.2B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
128K tokens
CAPABILITIES
chat, tool_use
RELEASE DATE
2024-07-18
PROVIDER
Mistral AI
FAMILY
mistral
▸ VRAM REQUIREMENTS
QUANT
BPW
VRAM
QUALITY
Q3_K_M
4
6.6 GB
88%
Q3_K_L
4.3
7.0 GB
90%
IQ4_XS
4.46
7.3 GB
92%
Q4_K_S
4.67
7.6 GB
93%
Q4_K_M
4.89
7.9 GB
94%
Q5_K_S
5.57
9.0 GB
96%
Q5_K_M
5.7
9.2 GB
96%
Q6_K
6.56
10.5 GB
97%
Q8_0
8.5
13.5 GB
100%
FP16
16
24.9 GB
100%
§ 01BENCHMARK SCORES
MMLU-PRO28.0
MATH12.7
IFEval63.8
BBH29.7
GPQA5.4
MUSR8.5
GPQA Diamond8.7
§ 02RUN COMMAND
Run Mistral-Nemo 12.2B locally with Ollama — needs 7.9 GB VRAM at Q4_K_M: