FitMyLLM
Moonshot AI/Mixture of Experts

MKimi K3

Kimi K3 — Moonshot's 2.8T MoE with 104B active, native multimodality and a 1M context. 69 of its 93 layers are Kimi Delta Attention and keep no growing KV cache; the other 24 use Gated MLA.

chatcodingreasoningmultilingualvisionmathagentictool_use
2779.93B
Parameters (104B active)
1024K
Context length
6
Benchmarks
17
Quantizations
2.1M
HF downloads
Architecture
MoE
Released
2026-07-16
Layers
93
KV Heads
96
Head Dim
74
Family
kimi

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ2_XXS2.38
827.8 GB
827.5 + 0.3 KV
low
IQ2_M2.93
1018.9 GB
1018.6 + 0.3 KV
low
Q2_K3.16
1098.9 GB
1098.6 + 0.3 KV
low
IQ3_XXS3.25
1130.1 GB
1129.8 + 0.3 KV
low
IQ3_XS3.5
1217.0 GB
1216.7 + 0.3 KV
low
Q3_K_S3.64
1265.7 GB
1265.4 + 0.3 KV
low
IQ3_M3.76
1307.4 GB
1307.1 + 0.3 KV
low
Q3_K_M4
1390.8 GB
1390.5 + 0.3 KV
low
Q3_K_L4.3
1495.0 GB
1494.7 + 0.3 KV
moderate
IQ4_XS4.46
1550.6 GB
1550.3 + 0.3 KV
moderate
Q4_K_S4.67
1623.6 GB
1623.3 + 0.3 KV
moderate
Q4_K_M4.89
1700.0 GB
1699.7 + 0.3 KV
good
Q5_K_S5.57
1936.3 GB
1936.0 + 0.3 KV
good
Q5_K_M5.7
1981.5 GB
1981.2 + 0.3 KV
good
Q6_K6.56
2280.3 GB
2280.0 + 0.3 KV
excellent
Q8_08.5
2954.5 GB
2954.2 + 0.3 KV
lossless
FP1616
5560.7 GB
5560.3 + 0.3 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Too big for a single GPU — plan a multi-GPU deployment
Even the lightest quant needs ~828 GB. Size GPUs, replicas, TCO and scaling for a production setup. Open in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN KIMI K3 NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (6)

GPQA Diamond93.5
AA Long Context82.7
AA Coding76.2
AA Intelligence59.7
SciCode58.7
HLE46.9

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run kimi:2780b-q4_K_M

Tag may need adjustment — check ollama.com/library/kimi for available tags.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

Find the best GPU for Kimi K3

Build Hardware for Kimi K3
▸ SPEC SHEET

Kimi K32779.93B MoE.

▸ SPECIFICATIONS
PARAMETERS
2779.93B (104B active)
ARCHITECTURE
Mixture of Experts
CONTEXT LENGTH
1024K tokens
CAPABILITIES
chat, coding, reasoning, multilingual, vision, math, agentic, tool_use
RELEASE DATE
2026-07-16
PROVIDER
Moonshot AI
FAMILY
kimi
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ2_XXS2.38827.5 GB65%
IQ2_M2.931018.6 GB75%
Q2_K3.161098.6 GB78%
IQ3_XXS3.251129.8 GB82%
IQ3_XS3.51216.7 GB84%
Q3_K_S3.641265.4 GB85%
IQ3_M3.761307.1 GB86%
Q3_K_M41390.5 GB88%
Q3_K_L4.31494.7 GB90%
IQ4_XS4.461550.3 GB92%
Q4_K_S4.671623.3 GB93%
Q4_K_M4.891699.7 GB94%
Q5_K_S5.571936.0 GB96%
Q5_K_M5.71981.2 GB96%
Q6_K6.562280.0 GB97%
Q8_08.52954.2 GB100%
FP16165560.3 GB100%
§ 01BENCHMARK SCORES
GPQA Diamond93.5
HLE46.9
AA Intelligence59.7
AA Coding76.2
aa_scicode58.7
aa_lcr82.7