FitMyLLM
DeepSeek/Mixture of Experts

DeepSeekDeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 — the July re-post-train of V4 Flash, same 284B / 13B-active architecture with large gains on agentic and coding work. The checkpoint weighs 304B because it ships a DSpark speculative-decoding draft module on top of the base model.

chatcodingreasoningmultilingualmathagentictool_use
304.18B
Parameters (13B active)
1024K
Context length
6
Benchmarks
17
Quantizations
1.9M
HF downloads
Architecture
MoE
Released
2026-07-31
Layers
43
KV Heads
1
Head Dim
512
Family
deepseek

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ2_XXS2.38
92.0 GB
91.0 + 1.0 KV
low
IQ2_M2.93
112.9 GB
111.9 + 1.0 KV
low
Q2_K3.16
121.6 GB
120.6 + 1.0 KV
low
IQ3_XXS3.25
125.1 GB
124.1 + 1.0 KV
low
IQ3_XS3.5
134.6 GB
133.6 + 1.0 KV
low
Q3_K_S3.64
139.9 GB
138.9 + 1.0 KV
low
IQ3_M3.76
144.5 GB
143.5 + 1.0 KV
low
Q3_K_M4
153.6 GB
152.6 + 1.0 KV
low
Q3_K_L4.3
165.0 GB
164.0 + 1.0 KV
moderate
IQ4_XS4.46
171.1 GB
170.1 + 1.0 KV
moderate
Q4_K_S4.67
179.1 GB
178.1 + 1.0 KV
moderate
Q4_K_M4.89
187.4 GB
186.4 + 1.0 KV
good
Q5_K_S5.57
213.3 GB
212.3 + 1.0 KV
good
Q5_K_M5.7
218.2 GB
217.2 + 1.0 KV
good
Q6_K6.56
250.9 GB
249.9 + 1.0 KV
excellent
Q8_08.5
324.7 GB
323.7 + 1.0 KV
lossless
FP1616
609.9 GB
608.8 + 1.0 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Too big for a single GPU — plan a multi-GPU deployment
Even the lightest quant needs ~92 GB. Size GPUs, replicas, TCO and scaling for a production setup. Open in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN DEEPSEEK V4 FLASH 0731 NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (6)

GPQA Diamond90.8
AA Long Context74.3
AA Coding69.1
AA Intelligence51.8
SciCode49.9
HLE38.6

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run deepseek:304b-q4_K_M

Tag may need adjustment — check ollama.com/library/deepseek for available tags.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

GPUs that can run this model

At Q4_K_M quantization. Sorted by minimum VRAM.

AMD Instinct MI300X
192 GB VRAM • 5300 GB/s
AMD
$15000
Apple M2 Ultra (192GB)
192 GB VRAM • 800 GB/s
APPLE
$5499
Apple M3 Ultra (192GB)
192 GB VRAM • 800 GB/s
APPLE
$6999
Apple M4 Ultra (192GB)
192 GB VRAM • 1092 GB/s
APPLE
$7499
AMD Radeon Instinct MI300A
192 GB VRAM • 10300 GB/s
AMD
$12000
AMD Radeon Instinct MI300X
192 GB VRAM • 10300 GB/s
AMD
$15000
AMD Radeon Instinct MI308X
192 GB VRAM • 10300 GB/s
AMD
$12000
Apple M5 Ultra (192GB)
192 GB VRAM • 1228 GB/s
APPLE
AMD Radeon Instinct MI325X
288 GB VRAM • 10300 GB/s
AMD
$20000
AMD Radeon Instinct MI350X
288 GB VRAM • 8190 GB/s
AMD
$25000
AMD Radeon Instinct MI355X
288 GB VRAM • 8190 GB/s
AMD
$30000
Apple M4 Ultra (384GB)
384 GB VRAM • 1092 GB/s
APPLE
$9999
Apple M5 Ultra (384GB)
384 GB VRAM • 1228 GB/s
APPLE

Find the best GPU for DeepSeek V4 Flash 0731

Build Hardware for DeepSeek V4 Flash 0731
▸ SPEC SHEET

DeepSeek V4 Flash 0731304.18B MoE.

▸ SPECIFICATIONS
PARAMETERS
304.18B (13B active)
ARCHITECTURE
Mixture of Experts
CONTEXT LENGTH
1024K tokens
CAPABILITIES
chat, coding, reasoning, multilingual, math, agentic, tool_use
RELEASE DATE
2026-07-31
PROVIDER
DeepSeek
FAMILY
deepseek
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ2_XXS2.3891.0 GB65%
IQ2_M2.93111.9 GB75%
Q2_K3.16120.6 GB78%
IQ3_XXS3.25124.1 GB82%
IQ3_XS3.5133.6 GB84%
Q3_K_S3.64138.9 GB85%
IQ3_M3.76143.5 GB86%
Q3_K_M4152.6 GB88%
Q3_K_L4.3164.0 GB90%
IQ4_XS4.46170.1 GB92%
Q4_K_S4.67178.1 GB93%
Q4_K_M4.89186.4 GB94%
Q5_K_S5.57212.3 GB96%
Q5_K_M5.7217.2 GB96%
Q6_K6.56249.9 GB97%
Q8_08.5323.7 GB100%
FP1616608.8 GB100%
§ 01BENCHMARK SCORES
GPQA Diamond90.8
HLE38.6
AA Intelligence51.8
AA Coding69.1
aa_scicode49.9
aa_lcr74.3
§ 03COMPATIBLE GPUs
13 @ Q4_K_M