FitMyLLM
StepFun/Mixture of Experts

SStep 3.7 Flash

Step 3.7 Flash — StepFun's 201B MoE with about 11B active and a built-in vision encoder, built for high-frequency production workloads.

chatcodingreasoningmultilingualvisionmath
201.37B
Parameters (11B active)
256K
Context length
9
Benchmarks
17
Quantizations
118K
HF downloads
Architecture
MoE
Released
2026-05-29
Layers
45
KV Heads
8
Head Dim
128
Family
other

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ2_XXS2.38
61.0 GB
60.4 + 0.6 KV
low
IQ2_M2.93
74.8 GB
74.2 + 0.6 KV
low
Q2_K3.16
80.6 GB
80.0 + 0.6 KV
low
IQ3_XXS3.25
82.9 GB
82.3 + 0.6 KV
low
IQ3_XS3.5
89.2 GB
88.6 + 0.6 KV
low
Q3_K_S3.64
92.7 GB
92.1 + 0.6 KV
low
IQ3_M3.76
95.7 GB
95.1 + 0.6 KV
low
Q3_K_M4
101.8 GB
101.2 + 0.6 KV
low
Q3_K_L4.3
109.3 GB
108.7 + 0.6 KV
moderate
IQ4_XS4.46
113.3 GB
112.8 + 0.6 KV
moderate
Q4_K_S4.67
118.6 GB
118.0 + 0.6 KV
moderate
Q4_K_M4.89
124.2 GB
123.6 + 0.6 KV
good
Q5_K_S5.57
141.3 GB
140.7 + 0.6 KV
good
Q5_K_M5.7
144.5 GB
144.0 + 0.6 KV
good
Q6_K6.56
166.2 GB
165.6 + 0.6 KV
excellent
Q8_08.5
215.0 GB
214.4 + 0.6 KV
lossless
FP1616
403.8 GB
403.2 + 0.6 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Too big for a single GPU — plan a multi-GPU deployment
Even the lightest quant needs ~61 GB. Size GPUs, replicas, TCO and scaling for a production setup. Open in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN STEP 3.7 FLASH NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (9)

τ²-Bench98.5
GPQA Diamond80.9
AA Long Context69.7
IFBench67.3
SciCode40.0
AA Coding39.6
Terminal-Bench35.6
AA Intelligence30.9
HLE21.4

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run other:201b-q4_K_M

Tag may need adjustment — check ollama.com/library/other for available tags.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

GPUs that can run this model

At Q4_K_M quantization. Sorted by minimum VRAM.

Apple M4 Max (128GB)
128 GB VRAM • 546 GB/s
APPLE
$3999
AMD Instinct MI250X
128 GB VRAM • 3277 GB/s
AMD
$10000
Apple M1 Ultra (128GB)
128 GB VRAM • 800 GB/s
APPLE
$4999
Apple M2 Ultra (128GB)
128 GB VRAM • 800 GB/s
APPLE
$3999
AMD Radeon Instinct MI250
128 GB VRAM • 3280 GB/s
AMD
$12000
AMD Radeon Instinct MI250X
128 GB VRAM • 3280 GB/s
AMD
$15000
AMD Radeon Instinct MI300
128 GB VRAM • 6550 GB/s
AMD
$12000
Intel Data Center GPU Max 1550
128 GB VRAM • 3280 GB/s
INTEL
Intel Data Center GPU Max Subsystem
128 GB VRAM • 3210 GB/s
INTEL
NVIDIA GB10
128 GB VRAM • 273 GB/s
NVIDIA
NVIDIA Jetson T5000
128 GB VRAM • 273 GB/s
NVIDIA
Apple M5 Max (128GB)
128 GB VRAM • 614 GB/s
APPLE
NVIDIA H200 SXM 141GB
140 GB VRAM • 4800 GB/s
NVIDIA
$30000
NVIDIA H200 NVL
141 GB VRAM • 4890 GB/s
NVIDIA
$35000
NVIDIA H200 SXM 141 GB
141 GB VRAM • 4890 GB/s
NVIDIA
$30000
NVIDIA B300
144 GB VRAM • 4100 GB/s
NVIDIA
$35000
AMD Instinct MI300X
192 GB VRAM • 5300 GB/s
AMD
$15000
Apple M2 Ultra (192GB)
192 GB VRAM • 800 GB/s
APPLE
$5499
Apple M3 Ultra (192GB)
192 GB VRAM • 800 GB/s
APPLE
$6999
Apple M4 Ultra (192GB)
192 GB VRAM • 1092 GB/s
APPLE
$7499
AMD Radeon Instinct MI300A
192 GB VRAM • 10300 GB/s
AMD
$12000
AMD Radeon Instinct MI300X
192 GB VRAM • 10300 GB/s
AMD
$15000
AMD Radeon Instinct MI308X
192 GB VRAM • 10300 GB/s
AMD
$12000
Apple M5 Ultra (192GB)
192 GB VRAM • 1228 GB/s
APPLE
AMD Radeon Instinct MI325X
288 GB VRAM • 10300 GB/s
AMD
$20000
AMD Radeon Instinct MI350X
288 GB VRAM • 8190 GB/s
AMD
$25000
AMD Radeon Instinct MI355X
288 GB VRAM • 8190 GB/s
AMD
$30000
Apple M4 Ultra (384GB)
384 GB VRAM • 1092 GB/s
APPLE
$9999
Apple M5 Ultra (384GB)
384 GB VRAM • 1228 GB/s
APPLE

Find the best GPU for Step 3.7 Flash

Build Hardware for Step 3.7 Flash
▸ SPEC SHEET

Step 3.7 Flash201.37B MoE.

▸ SPECIFICATIONS
PARAMETERS
201.37B (11B active)
ARCHITECTURE
Mixture of Experts
CONTEXT LENGTH
256K tokens
CAPABILITIES
chat, coding, reasoning, multilingual, vision, math
RELEASE DATE
2026-05-29
PROVIDER
StepFun
FAMILY
other
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ2_XXS2.3860.4 GB65%
IQ2_M2.9374.2 GB75%
Q2_K3.1680.0 GB78%
IQ3_XXS3.2582.3 GB82%
IQ3_XS3.588.6 GB84%
Q3_K_S3.6492.1 GB85%
IQ3_M3.7695.1 GB86%
Q3_K_M4101.2 GB88%
Q3_K_L4.3108.7 GB90%
IQ4_XS4.46112.8 GB92%
Q4_K_S4.67118.0 GB93%
Q4_K_M4.89123.6 GB94%
Q5_K_S5.57140.7 GB96%
Q5_K_M5.7144.0 GB96%
Q6_K6.56165.6 GB97%
Q8_08.5214.4 GB100%
FP1616403.2 GB100%
§ 01BENCHMARK SCORES
GPQA Diamond80.9
HLE21.4
AA Intelligence30.9
AA Coding39.6
aa_ifbench67.3
aa_terminal_bench35.6
aa_tau298.5
aa_scicode40.0
aa_lcr69.7
§ 03COMPATIBLE GPUs
29 @ Q4_K_M