FITMYLLM · JULY 28, 2026 · OPEN METHODOLOGY · COMMUNITY BENCHMARKS
LIVE · UPDATED DAILY
35B
Parameters (3B active)
Quantization Options Select your GPU for speed estimates Context length: 4K 8K 16K 32K 64K 128K 256K
Quant Bits VRAM @ 16K Quality IQ3_XXS 3.25 15.6 GB
14.7 + 0.9 KV
low IQ3_XS 3.5 16.7 GB
15.8 + 0.9 KV
low Q3_K_S 3.64 17.4 GB
16.4 + 0.9 KV
low IQ3_M 3.76 17.9 GB
16.9 + 0.9 KV
low Q3_K_M 4 18.9 GB
18.0 + 0.9 KV
low Q3_K_L 4.3 20.2 GB
19.3 + 0.9 KV
moderate IQ4_XS 4.46 20.9 GB
20.0 + 0.9 KV
moderate Q4_K_S 4.67 21.9 GB
20.9 + 0.9 KV
moderate Q4_K_M 4.89 22.8 GB
21.9 + 0.9 KV
good Q5_K_S 5.57 25.8 GB
24.9 + 0.9 KV
good Q5_K_M 5.7 26.4 GB
25.4 + 0.9 KV
good Q6_K 6.56 30.1 GB
29.2 + 0.9 KV
excellent Q8_0 8.5 38.6 GB
37.7 + 0.9 KV
lossless FP16 16 71.4 GB
70.5 + 0.9 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Deploying for a team or in production? Size GPUs, cost & scaling in Enterprise → ▸ READY TO RUN THIS? RENT BY THE HOUR
RENT A GPU AND RUN QWEN 3.5 35B A3B NOW
Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.
Community Ratings Chat Coding Reasoning Creative Vision Roleplay Agentic
Loading ratings...
Run this model IQ3_XXS — 14.7 GB VRAM IQ3_XS — 15.8 GB VRAM Q3_K_S — 16.4 GB VRAM IQ3_M — 16.9 GB VRAM Q3_K_M — 18.0 GB VRAM Q3_K_L — 19.3 GB VRAM IQ4_XS — 20.0 GB VRAM Q4_K_S — 20.9 GB VRAM Q4_K_M — 21.9 GB VRAM Q5_K_S — 24.9 GB VRAM Q5_K_M — 25.4 GB VRAM Q6_K — 29.2 GB VRAM Q8_0 — 37.7 GB VRAM FP16 — 70.5 GB VRAM
Ollama llama.cpp vLLM LM Studio KoboldCpp Jan Docker
▸ Easiest way to get started · Beginners
DOCS ↗ curl -fsSL https://ollama.com/install.sh | shCOPY
$ ollama run qwen3.5:35b-a3b-q4_K_MCOPY
Downloads and runs automatically. Add --verbose for speed stats.
▸ SETUP GUIDE >_
Auto-setup with fitmyllm CLI Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.
Auto-detect GPU Live tok/s in chat Speed benchmarks 9 inference engines
GPUs that can run this model At Q4_K_M quantization. Sorted by minimum VRAM.
Find the best GPU for Qwen 3.5 35B A3B
Build Hardware for Qwen 3.5 35B A3B Qwen 3.5 35B A3B — MoE with multimodal, math, vision. Runs at 3B speed.
Read full model card ▸ COLOPHON FITMYLLM · INDEPENDENT · DATA-DRIVEN
FITMYLLM · EST. 2025 · © 2026
RECOMMENDATIONS FROM PUBLISHED MATH, CORRECTED BY THE COMMUNITY — 30.
▸ SPEC SHEET
Qwen 3.5 35B A3B — 35B MoE. ▸ SPECIFICATIONS
PARAMETERS 35B (3B active)
ARCHITECTURE Mixture of Experts
CONTEXT LENGTH 256K tokens
CAPABILITIES chat, coding, reasoning, multilingual, vision, math
RELEASE DATE 2026-02-01
PROVIDER Alibaba
FAMILY qwen ▸ VRAM REQUIREMENTS
QUANT BPW VRAM QUALITY IQ3_XXS 3.25 14.7 GB 82% IQ3_XS 3.5 15.8 GB 84% Q3_K_S 3.64 16.4 GB 85% IQ3_M 3.76 16.9 GB 86% Q3_K_M 4 18.0 GB 88% Q3_K_L 4.3 19.3 GB 90% IQ4_XS 4.46 20.0 GB 92% Q4_K_S 4.67 20.9 GB 93% Q4_K_M 4.89 21.9 GB 94% Q5_K_S 5.57 24.9 GB 96% Q5_K_M 5.7 25.4 GB 96% Q6_K 6.56 29.2 GB 97% Q8_0 8.5 37.7 GB 100% FP16 16 70.5 GB 100%
§ 01 BENCHMARK SCORES
MMLU-PRO 85.3
MATH 59.7
IFEval 91.9
BBH 58.3
MMMU 75.1
GPQA 15.2
MUSR 19.1
BigCodeBench 32.3
MMBench 91.5
Arena Elo 1485.0
GPQA Diamond 81.9
HLE 12.8
AA Intelligence 30.7
AA Coding 16.8
aa_ifbench 72.5
aa_terminal_bench 26.5
aa_tau2 89.2
aa_scicode 37.7
aa_lcr 62.7
§ 02 RUN COMMAND
Run Qwen 3.5 35B A3B locally with Ollama — needs 21.9 GB VRAM at Q4_K_M:
$ ollama run qwen3.5:35b
§ 03 COMPATIBLE GPUs
30 @ Q4_K_M Feedback