Nemotron 3 Super 120B-A12B
NVIDIA's 120B-class hybrid LatentMoE activates about 12B parameters per token and interleaves Mamba-2, attention, and expert blocks. The official checkpoint supports configurable reasoning and contexts up to one million tokens.
NVIDIAMixture of expertsagenticreasoning17 machines can run it
Total parameters
123.611B
Official spec sets memory need
Active parameters
12B
Official spec sets decode speed
Context
1M
Official spec tokens
4-bit weights
67 GB
Calculated before KV cache
This is a sparse mixture of experts. All 123.611B of weights must be available to the runtime, but only 12B are read per generated token. Ordinary runtimes keep the full quantized model in memory; a runtime-specific SSD-streaming artifact can retain a smaller working set and fetch expert data on demand, trading speed for capacity.
Model specification
Publisher
NVIDIA Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
123.611B Official spec
Active parameters
12B per token Official spec
Context length
1,048,576 tokens Official spec
Attention ⓘ
Hybrid Mamba-2 · 8 attention + 40 Mamba blocks Official spec
Licence
NVIDIA Nemotron Open Model License Official spec source
Specification confidence ⓘ
High · verified 2026-09-21 Official spec source
Released
11 Mar 2026 Official spec
Official source
Quantizations and memory
| Quantization | Format | Bits/weight | Resident | Download | Quality kept |
|---|---|---|---|---|---|
MLX 4-bit | mlx | 4.5 | 67.3 GB | 67.3 GB | 98.0% |
Q4_K_Mdefault | gguf | 4.85 | 71.9 GB | 71.9 GB | 98.5% |
Q5_K_M | gguf | 5.69 | 84.3 GB | 84.3 GB | 99.3% |
Q8_0 | gguf | 8.5 | 124.8 GB | 124.8 GB | 99.9% |
BF16 | safetensors | 16 | 230.2 GB | 230.2 GB | 100.0% |
Generic weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead. Runtime-specific artifacts use their published resident and download footprints; streamed models can therefore require much more disk than memory. Quality retention is an assumption, not a measured evaluation.
Coding & quality benchmarksCompare coding results →
No published benchmark results have been imported for this model yet. This is missing evidence, not a score of zero.
We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
~2.4–3.5 t/s · $1,531
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~17–24.4 t/s · $3,873
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~48.5–69.8 t/s · $4,323
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per purchase-price unit
~95.9–138 t/s · $5,945
Highest decode tokens/second per 1,000 units of the displayed purchase currency. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Nemotron 3 Super 120B-A12B0 measured, 17 estimated
17 of 17 rows
| Hardware↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Price↕ | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|---|
| NVIDIA DGX H200 (8x H200, 1,128 GB) NVIDIA · 1,128 GB · 38,400 GB/s | Q8_0recommended | 132.5 GB | ~421–605 t/s | ~69470–144290 t/s | 1024K | $359,430 | Comfortable | Estimated |
| Quad RTX 5090 workstation (4x 32 GB) NVIDIA · 128 GB · 7,168 GB/s | Q5_K_Mrecommended | 89.3 GB | ~178–256 t/s | ~2600–5400 t/s | 1024K | $12,611 | Comfortable | Estimated |
| Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB) Lenovo · 384 GB · 7,168 GB/s | Q8_0recommended | 130.9 GB | ~117–169 t/s | ~3330–6930 t/s | 1024K | $58,553 | Comfortable | Estimated |
| RTX PRO 6000 Blackwell workstation (96 GB) NVIDIA · 96 GB · 1,792 GB/s | Q4_K_Mrecommended | 75.3 GB | ~113–163 t/s | ~845–1750 t/s | 1024K | $10,359 | Fits | Estimated |
| RTX PRO 6000 Max-Q workstation (96 GB, 300 W) NVIDIA · 96 GB · 1,792 GB/s | Q4_K_Mrecommended | 75.3 GB | ~113–163 t/s | ~692–1440 t/s | 1024K | $10,809 | Fits | Estimated |
| Mac Studio M5 Ultra 96 GB Apple · 96 GB · 1,200 GB/s | MLX 4-bitrecommended | 70.6 GB | ~95.9–138 t/s | ~603–1250 t/s | 1024K | $5,945 | Fits | Estimated |
| Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB) Dell · 192 GB · 3,584 GB/s | Q8_0recommended | 130.1 GB | ~84.5–122 t/s | ~1670–3460 t/s | 1024K | $36,663 | Comfortable | Estimated |
| NVIDIA DGX Station GB300 (748 GB) NVIDIA · 748 GB · 7,100 GB/s | Q8_0recommended | 129.7 GB | ~54.4–78.2 t/s | ~9280–19280 t/s | 1024K | $90,082 | Comfortable | Estimated |
| Mac Studio M5 Ultra 512 GB Apple · 512 GB · 1,200 GB/s | Q8_0recommended | 129.7 GB | ~53.7–77.3 t/s | ~754–1570 t/s | 1024K | $11,350 | Comfortable | Estimated |
| Mac Studio M5 Ultra 256 GB Apple · 256 GB · 1,200 GB/s | Q8_0recommended | 129.7 GB | ~53.7–77.3 t/s | ~754–1570 t/s | 1024K | $8,647 | Comfortable | Estimated |
| Mac Studio M5 Max 128 GB Apple · 128 GB · 614 GB/s | Q4_K_Mrecommended | 75.3 GB | ~48.5–69.8 t/s | ~377–783 t/s | 1024K | $4,323 | Comfortable | Estimated |
| Mac Studio M3 Ultra 512 GB Apple · 512 GB · 819 GB/s | Q8_0recommended | 129.7 GB | ~30.5–43.9 t/s | ~182–377 t/s | 1024K | $8,287 | Comfortable | Estimated |
| Mac Studio M3 Ultra 256 GB Apple · 256 GB · 819 GB/s | Q8_0recommended | 129.7 GB | ~30.5–43.9 t/s | ~182–377 t/s | 1024K | $5,674 | Comfortable | Estimated |
| NVIDIA DGX Spark 128 GB NVIDIA · 128 GB · 273 GB/s | Q4_K_Mrecommended | 75.3 GB | ~17–24.4 t/s | ~324–674 t/s | 1024K | $3,873 | Comfortable | Estimated |
| Framework Desktop (Ryzen AI Max+ 395, 128 GB) AMD · 128 GB · 256 GB/s | Q4_K_Mrecommended | 75.3 GB | ~11.6–16.7 t/s | ~89.6–186 t/s | 1024K | $2,161 | Comfortable | Estimated |
| GMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB) AMD · 128 GB · 256 GB/s | Q4_K_Mrecommended | 75.3 GB | ~10.1–14.5 t/s | ~89.6–186 t/s | 1024K | $1,711 | Comfortable | Estimated |
| CPU-only workstation (Ryzen 9950X, 192 GB DDR5) Generic · 192 GB · 90 GB/s | Q8_0recommended | 129.7 GB | ~2.4–3.5 t/s | ~10.8–22.4 t/s | 1024K | $1,531 | Comfortable | Estimated |
Submit a benchmarkContributions are reviewed before publication.
Similar models
Qwen3.8 Flash Next
Alibaba Qwen · 180B (6B active) · Qwen Community 1.0
Nemotron 3.5 Lightning 30B-A3B
NVIDIA · 31.578B (3B active) · OpenMDW 1.1
Step 3.7 Flash 198B-A11B
StepFun · 201.365B (11B active) · Apache 2.0
MiMo-V2.6 Flash 309B-A15B
XiaomiMiMo · 309B (15B active) · MIT
MiniMax M3
MiniMax · 427.04B (23B active) · Apache 2.0
Qwen3.5 122B-A10B
Alibaba Qwen · 125.086B (10B active) · Apache 2.0