The new xLAM-2 series, built on our most advanced data synthesis, processing, and training pipelines, marks a significant leap in multi-turn conversation and tool usage.
tool_useagentic
70B
Parameters
128K
Context length
1
Benchmarks
16
Quantizations
0
Architecture
Dense
Released
2025-03-12
Layers
80
KV Heads
8
Head Dim
128
Family
other
Quantization Options
Context length:
Quant
Bits
VRAM @ 16K
Quality
IQ2_M
2.93
29.9 GB
26.1 + 3.8 KV
low
Q2_K
3.16
31.9 GB
28.1 + 3.8 KV
low
IQ3_XXS
3.25
32.7 GB
28.9 + 3.8 KV
low
IQ3_XS
3.5
34.9 GB
31.1 + 3.8 KV
low
Q3_K_S
3.64
36.1 GB
32.3 + 3.8 KV
low
IQ3_M
3.76
37.1 GB
33.4 + 3.8 KV
low
Q3_K_M
4
39.2 GB
35.5 + 3.8 KV
low
Q3_K_L
4.3
41.9 GB
38.1 + 3.8 KV
moderate
IQ4_XS
4.46
43.3 GB
39.5 + 3.8 KV
moderate
Q4_K_S
4.67
45.1 GB
41.4 + 3.8 KV
moderate
Q4_K_M
4.89
47.0 GB
43.3 + 3.8 KV
good
Q5_K_S
5.57
53.0 GB
49.2 + 3.8 KV
good
Q5_K_M
5.7
54.1 GB
50.4 + 3.8 KV
good
Q6_K
6.56
61.6 GB
57.9 + 3.8 KV
excellent
Q8_0
8.5
78.6 GB
74.9 + 3.8 KV
lossless
FP16
16
144.2 GB
140.5 + 3.8 KV
lossless
Select your GPU above to see speed estimates and compatibility for each quantization.
Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.