FitMyLLM
exaone/Dense

EEXAONE Deep 32B

We introduce EXAONE Deep, which exhibits superior capabilities in various reasoning tasks including math and coding benchmarks, ranging from 2.4B to 32B parameters developed and released by LG AI Research.

reasoningmathcoding
32B
Parameters
32K
Context length
10
Benchmarks
14
Quantizations
Architecture
Dense
Released
2025-03-16
Layers
64
KV Heads
8
Head Dim
128
Family
exaone

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ3_XXS3.25
16.5 GB
13.5 + 3.0 KV
low
IQ3_XS3.5
17.5 GB
14.5 + 3.0 KV
low
Q3_K_S3.64
18.0 GB
15.0 + 3.0 KV
low
IQ3_M3.76
18.5 GB
15.5 + 3.0 KV
low
Q3_K_M4
19.5 GB
16.5 + 3.0 KV
low
Q3_K_L4.3
20.7 GB
17.7 + 3.0 KV
moderate
IQ4_XS4.46
21.3 GB
18.3 + 3.0 KV
moderate
Q4_K_S4.67
22.2 GB
19.2 + 3.0 KV
moderate
Q4_K_M4.89
23.0 GB
20.0 + 3.0 KV
good
Q5_K_S5.57
25.8 GB
22.8 + 3.0 KV
good
Q5_K_M5.7
26.3 GB
23.3 + 3.0 KV
good
Q6_K6.56
29.7 GB
26.7 + 3.0 KV
excellent
Q8_08.5
37.5 GB
34.5 + 3.0 KV
lossless
FP1616
67.5 GB
64.5 + 3.0 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Deploying for a team or in production? Size GPUs, cost & scaling in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN EXAONE DEEP 32B NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (10)

MATH-50095.7
IFEval83.9
AIME72.1
GPQA Diamond66.1
LiveCodeBench59.5
MATH51.3
MMLU-PRO40.4
BBH39.8
MUSR5.2
GPQA5.0

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run exaone-deep:32b-q4_K_M

Downloads and runs automatically. Add --verbose for speed stats.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

GPUs that can run this model

At Q4_K_M quantization. Sorted by minimum VRAM.

Apple M4 Pro (24GB)
24 GB VRAM • 273 GB/s
APPLE
$1399
NVIDIA L4 24GB
24 GB VRAM • 300 GB/s
NVIDIA
$2500
Apple M2 (24GB)
24 GB VRAM • 100 GB/s
APPLE
$999
Apple M3 (24GB)
24 GB VRAM • 100 GB/s
APPLE
$999
Apple M4 (24GB)
24 GB VRAM • 120 GB/s
APPLE
$699
NVIDIA Tesla M40 24 GB
24 GB VRAM • 288 GB/s
NVIDIA
NVIDIA Tesla P10
24 GB VRAM • 694 GB/s
NVIDIA
NVIDIA Tesla P40
24 GB VRAM • 347 GB/s
NVIDIA
NVIDIA RTX A5000
24 GB VRAM • 768 GB/s
NVIDIA
$2500
NVIDIA L40 CNX
24 GB VRAM • 864 GB/s
NVIDIA
$5000
NVIDIA L40G
24 GB VRAM • 864 GB/s
NVIDIA
$5000

Find the best GPU for EXAONE Deep 32B

Build Hardware for EXAONE Deep 32B
▸ SPEC SHEET

EXAONE Deep 32B32B Dense.

▸ SPECIFICATIONS
PARAMETERS
32B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
32K tokens
CAPABILITIES
reasoning, math, coding
RELEASE DATE
2025-03-16
FAMILY
exaone
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ3_XXS3.2513.5 GB82%
IQ3_XS3.514.5 GB84%
Q3_K_S3.6415.0 GB85%
IQ3_M3.7615.5 GB86%
Q3_K_M416.5 GB88%
Q3_K_L4.317.7 GB90%
IQ4_XS4.4618.3 GB92%
Q4_K_S4.6719.2 GB93%
Q4_K_M4.8920.0 GB94%
Q5_K_S5.5722.8 GB96%
Q5_K_M5.723.3 GB96%
Q6_K6.5626.7 GB97%
Q8_08.534.5 GB100%
FP161664.5 GB100%
§ 01BENCHMARK SCORES
MMLU-PRO40.4
MATH51.3
IFEval83.9
BBH39.8
GPQA5.0
MUSR5.2
LiveCodeBench59.5
AIME72.1
MATH-50095.7
GPQA Diamond66.1
§ 02RUN COMMAND

Run EXAONE Deep 32B locally with Ollama — needs 20.0 GB VRAM at Q4_K_M:

$ollama run exaone-deep:32b
§ 03COMPATIBLE GPUs
30 @ Q4_K_M