HuatuoGPT-o1 is a medical LLM designed for advanced medical reasoning. It generates a complex thought process, reflecting and refining its reasoning, before providing a final response.
chatreasoning
70B
Parameters
8K
Context length
0
Benchmarks
16
Quantizations
0
Architecture
Dense
Released
2024-12-25
Layers
80
KV Heads
8
Head Dim
128
Family
other
Quantization Options
Context length:
Quant
Bits
VRAM @ 8K
Quality
IQ2_M
2.93
27.4 GB
26.1 + 1.3 KV
low
Q2_K
3.16
29.4 GB
28.1 + 1.3 KV
low
IQ3_XXS
3.25
30.2 GB
28.9 + 1.3 KV
low
IQ3_XS
3.5
32.4 GB
31.1 + 1.3 KV
low
Q3_K_S
3.64
33.6 GB
32.3 + 1.3 KV
low
IQ3_M
3.76
34.6 GB
33.4 + 1.3 KV
low
Q3_K_M
4
36.7 GB
35.5 + 1.3 KV
low
Q3_K_L
4.3
39.4 GB
38.1 + 1.3 KV
moderate
IQ4_XS
4.46
40.8 GB
39.5 + 1.3 KV
moderate
Q4_K_S
4.67
42.6 GB
41.4 + 1.3 KV
moderate
Q4_K_M
4.89
44.5 GB
43.3 + 1.3 KV
good
Q5_K_S
5.57
50.5 GB
49.2 + 1.3 KV
good
Q5_K_M
5.7
51.6 GB
50.4 + 1.3 KV
good
Q6_K
6.56
59.1 GB
57.9 + 1.3 KV
excellent
Q8_0
8.5
76.1 GB
74.9 + 1.3 KV
lossless
FP16
16
141.7 GB
140.5 + 1.3 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.