FitMyLLM
NVIDIA/Dense

NVIDIANVLM-D 72B

This family of models performs vision-language and text-only tasks including optical character recognition, multimodal reasoning, localization, common sense reasoning, world knowledge utilization, and coding.

chatvisionreasoning
79.38B
Parameters
32K
Context length
9
Benchmarks
16
Quantizations
0
Architecture
Dense
Released
2024-10-08
Layers
80
KV Heads
8
Head Dim
128
Family
nemotron

Quantization Options

Context length:
QuantBitsVRAM @ 16KQuality
IQ2_M2.93
33.3 GB
29.6 + 3.8 KV
low
Q2_K3.16
35.6 GB
31.8 + 3.8 KV
low
IQ3_XXS3.25
36.5 GB
32.7 + 3.8 KV
low
IQ3_XS3.5
39.0 GB
35.2 + 3.8 KV
low
Q3_K_S3.64
40.4 GB
36.6 + 3.8 KV
low
IQ3_M3.76
41.5 GB
37.8 + 3.8 KV
low
Q3_K_M4
43.9 GB
40.2 + 3.8 KV
low
Q3_K_L4.3
46.9 GB
43.2 + 3.8 KV
moderate
IQ4_XS4.46
48.5 GB
44.7 + 3.8 KV
moderate
Q4_K_S4.67
50.6 GB
46.8 + 3.8 KV
moderate
Q4_K_M4.89
52.8 GB
49.0 + 3.8 KV
good
Q5_K_S5.57
59.5 GB
55.8 + 3.8 KV
good
Q5_K_M5.7
60.8 GB
57.0 + 3.8 KV
good
Q6_K6.56
69.3 GB
65.6 + 3.8 KV
excellent
Q8_08.5
88.6 GB
84.8 + 3.8 KV
lossless
FP1616
163.0 GB
159.2 + 3.8 KV
lossless

Select your GPU above to see speed estimates and compatibility for each quantization.

Deploying for a team or in production? Size GPUs, cost & scaling in Enterprise →
READY TO RUN THIS?RENT BY THE HOUR

RENT A GPU AND RUN NVLM-D 72B NOW

Spin up an A100 / H100 / 4090 in ~60s. Pay by the second. Cancel anytime.

Community Ratings

Loading ratings...

Benchmarks (9)

HumanEval89.0
IFEval75.3
MMMU58.7
BBH55.8
MMLU-PRO48.1
BigCodeBench38.7
MATH33.3
GPQA20.8
MUSR18.4

Run this model

Easiest way to get started·Beginners
DOCS ↗
curl -fsSL https://ollama.com/install.sh | sh
$ollama run nemotron:79b-q4_K_M

Tag may need adjustment — check ollama.com/library/nemotron for available tags.

▸ SETUP GUIDE
>_

Auto-setup with fitmyllm CLI

Detects your GPU, recommends the best model, downloads it, and starts chatting — zero config. Benchmarks your speed and contributes anonymous data to improve predictions.

pip install fitmyllmthen run fitmyllmLearn more
Auto-detect GPULive tok/s in chatSpeed benchmarks9 inference engines

GPUs that can run this model

At Q4_K_M quantization. Sorted by minimum VRAM.

Apple M1 Ultra (64GB)
64 GB VRAM • 800 GB/s
APPLE
$2499
Apple M2 Ultra (64GB)
64 GB VRAM • 800 GB/s
APPLE
$2999
Apple M4 Max (64GB)
64 GB VRAM • 546 GB/s
APPLE
$2899
Apple M2 Max (64GB)
64 GB VRAM • 400 GB/s
APPLE
$2299
Apple M3 Max (64GB)
64 GB VRAM • 300 GB/s
APPLE
$2799
Apple M4 Pro (64GB)
64 GB VRAM • 273 GB/s
APPLE
$2599
AMD Radeon Instinct MI200
64 GB VRAM • 1640 GB/s
AMD
$10000
AMD Radeon Instinct MI210
64 GB VRAM • 1640 GB/s
AMD
$8000
NVIDIA H100 SXM5 64 GB
64 GB VRAM • 2020 GB/s
NVIDIA
$25000
NVIDIA Jetson AGX Orin 64 GB
64 GB VRAM • 205 GB/s
NVIDIA
NVIDIA Jetson T4000
64 GB VRAM • 273 GB/s
NVIDIA
Apple M5 Pro (64GB)
64 GB VRAM • 200 GB/s
APPLE
Apple M5 Max (64GB)
64 GB VRAM • 614 GB/s
APPLE
NVIDIA RTX PRO 5000 72 GB Blackwell
72 GB VRAM • 1340 GB/s
NVIDIA
$6999
NVIDIA H100 SXM5 80GB
80 GB VRAM • 3350 GB/s
NVIDIA
$25000
NVIDIA H100 PCIe 80GB
80 GB VRAM • 2000 GB/s
NVIDIA
$25000
NVIDIA A100 SXM 80GB
80 GB VRAM • 2039 GB/s
NVIDIA
$10000
NVIDIA A100 PCIe 80GB
80 GB VRAM • 1935 GB/s
NVIDIA
$10000
NVIDIA A100 SXM4 80 GB
80 GB VRAM • 2040 GB/s
NVIDIA
$15000
NVIDIA A100 PCIe 80 GB
80 GB VRAM • 1940 GB/s
NVIDIA
$10000
NVIDIA A100X
80 GB VRAM • 2040 GB/s
NVIDIA
NVIDIA H100 PCIe 80 GB
80 GB VRAM • 2040 GB/s
NVIDIA
$25000
NVIDIA H100 SXM5 80 GB
80 GB VRAM • 3360 GB/s
NVIDIA
$25000
NVIDIA H100 CNX
80 GB VRAM • 2040 GB/s
NVIDIA
$25000
NVIDIA A800 PCIe 80 GB
80 GB VRAM • 1940 GB/s
NVIDIA
NVIDIA A800 SXM4 80 GB
80 GB VRAM • 2040 GB/s
NVIDIA
NVIDIA H800 PCIe 80 GB
80 GB VRAM • 2040 GB/s
NVIDIA
NVIDIA H800 SXM5
80 GB VRAM • 3360 GB/s
NVIDIA
NVIDIA RTX 6000D
84 GB VRAM • 1570 GB/s
NVIDIA
$7500
NVIDIA B200
90 GB VRAM • 4100 GB/s
NVIDIA
$30000

Find the best GPU for NVLM-D 72B

Build Hardware for NVLM-D 72B
▸ SPEC SHEET

NVLM-D 72B79.38B Dense.

▸ SPECIFICATIONS
PARAMETERS
79.38B
ARCHITECTURE
Dense Transformer
CONTEXT LENGTH
32K tokens
CAPABILITIES
chat, vision, reasoning
RELEASE DATE
2024-10-08
PROVIDER
NVIDIA
FAMILY
nemotron
▸ VRAM REQUIREMENTS
QUANTBPWVRAMQUALITY
IQ2_M2.9329.6 GB75%
Q2_K3.1631.8 GB78%
IQ3_XXS3.2532.7 GB82%
IQ3_XS3.535.2 GB84%
Q3_K_S3.6436.6 GB85%
IQ3_M3.7637.8 GB86%
Q3_K_M440.2 GB88%
Q3_K_L4.343.2 GB90%
IQ4_XS4.4644.7 GB92%
Q4_K_S4.6746.8 GB93%
Q4_K_M4.8949.0 GB94%
Q5_K_S5.5755.8 GB96%
Q5_K_M5.757.0 GB96%
Q6_K6.5665.6 GB97%
Q8_08.584.8 GB100%
FP1616159.2 GB100%
§ 01BENCHMARK SCORES
HumanEval89.0
MMLU-PRO48.1
MATH33.3
IFEval75.3
BBH55.8
MMMU58.7
GPQA20.8
MUSR18.4
BigCodeBench38.7
§ 03COMPATIBLE GPUs
30 @ Q4_K_M