Nemotron 3 Super 120B-A12B on RTX PRO 6000 Blackwell workstation (96 GB)

123.611B (12B active) on 96 GB at 1,792 GB/s. Hardware details · Model details

Run locally →
Yes. Nemotron 3 Super 120B-A12B fits on RTX PRO 6000 Blackwell workstation (96 GB) at the recommended Q4_K_M configuration, requiring approximately 75.3 GB at 8K context. Practical context capacity is 1M. Expected decode for the recommended configuration is ~113–163 t/s. Estimated
CompatibilityFits
Calculated fit
Yes
Calculated Uses 80-90% of usable memory. Works, with limited headroom.
Recommended quantization
Q4_K_M
Calculated gguf
Memory required
75.3 GB
Calculated of 88.3 GB usable — 85%
Max practical context
1M
Calculated model supports 1M

Memory budget at 8K context

Model weights
71.9 GB Calculated
KV cache
0.2 GB Calculated
Runtime overhead
3.2 GB Estimated
Total required
75.3 GB Calculated
Headroom
13.1 GB Calculated
Fits with limited headroom (85% of usable memory). A longer context will not leave room for much else.
Discrete GPU: 96 GB of VRAM, of which we assume 92% is usable after driver and context overhead.
Mixture of experts: all 123.611B parameters must be resident in memory even though only ~12B are active per token. Memory follows total parameters; speed follows active parameters.
Hybrid cache: token-growing KV memory is charged only to 8 attention blocks; 40 Mamba-2 blocks use fixed-size convolution and recurrent state instead.
Performance

Estimated performance · recommended Q4_K_M

Decode, prefill and TTFT below are estimates for Q4_K_M. We do not have a comparable Q4_K_M measurement on this machine.

Estimated decode
~113–163 t/s
Estimated Q4_K_M; calculated range
Estimated prefill
~845–1750 t/s
Estimated Q4_K_M; calculated range
Estimated TTFT at 8K
~4.7–9.8 s
Estimated Q4_K_M; calculated range
Hardware load reference
660 W
Measured machine-level load; not this model run
How the estimate is calculated
  1. Decode: reading 12.0B active parameters at 4.85 bits/weight takes 5.46 ms at 1792 GB/s x 74% achieved efficiency.
  2. MoE routing penalty of 15% applied: expert gathers are less bandwidth-efficient than a dense sweep.
  3. Prefill: 232 TFLOPS (FP16) x 2 for native FP8 tensor cores x 0.67 calibrated against measured prefill on this platform x 32% assumed model-FLOPs utilisation, divided by 2 x 38.5B parameters per token.
  4. Prefill uses 38.5B effective parameters, not the 12B active in decode: a batch of hundreds of tokens routes across most of the expert pool.

  • MoE decode depends on how well the engine batches expert gathers; real results vary more than for dense models.
  • Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
  • This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.

Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.

Submit a benchmarkSubmissions stay pending until reviewed.
This model at other quantizations on this machine
QuantizationResidentDiskTotal neededUtilisationFitMax context
Q4_K_Mrecommended71.9 GB71.9 GB75.3 GB85%Fits1M
Q5_K_M84.3 GB84.3 GB88.1 GB100%Does not fit0
Q8_0124.8 GB124.8 GB129.7 GB147%Does not fit0
BF16230.2 GB230.2 GB238.4 GB270%Does not fit0
Cost of running Nemotron 3 Super 120B-A12B on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$134
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $34 more
Net cash at purchase
$10,359
Calculated VAT not reclaimable
Total over 5 years
$8,031
Calculated after tax, after resale
Cost per USD/1M tokens
$7.63
Calculated 17.5M tokens/month

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Purchase-price input: $10,359 in United States (federal), tax/VAT excluded · estimated · NL workstation integrator (converted catalogue estimate). This is the localized purchase input; the after-tax economic cost is calculated separately below.

Monthly breakdown

Depreciation
8,805 over 5 years, straight-line to a 1,554 residual
$146.75
Electricity
31.7 kWh/month at 0.140/kWh
$4.44
Cost of capital
4.0%/yr on 5,956 average capital employed
$19.85
Monthly cost before tax$171.04
Electricity tax shield
Running costs are deductible business expenses
−$0.93
First-year expensing
§179 (100.0%)
−$36.26
Monthly economic cost after tax$133.85

Three different numbers, deliberately

Cash cost
$10,359
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$146.75/month
$8,805 written down over 5 years to a $1,554 residual.
After-tax economic cost
$133.85/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$1,554 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
180 W Calculated
Load 660 W for 20% of powered hours, idle 60 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $2,175 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$19.85/month Assumption
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.
Efficiency metrics
Decode per $1,000 spent
13.4 t/s
Calculated purchase price only
Decode per $100/month
103.4 t/s
Calculated after-tax ownership cost
Tokens per joule
0.77
Calculated same as tokens/s per watt
USD per 1M output tokens
$7.63
Calculated at 20% utilisation
Local versus hosted APIs
Hosted modelUSD/1M outputBreak-evenAPI at your volumeVerdictComparison type
Qwen3.8 Flash
Alibaba Cloud
$0.4286.9M/mo$27API cheaperDifferent model
DeepSeek-V4.1 Flash
DeepSeek
$0.6074.4M/mo$32API cheaperDifferent model
DeepSeek-V4 Pro
DeepSeek
$1.9818.4M/mo$127About equalDifferent model
DeepSeek V4 Pro (Together)
Together AI
$4.406.3M/mo$372Local cheaperDifferent model
GLM-5.3
Z.ai
$4.408.6M/mo$274Local cheaperDifferent model
Claude Haiku 4.5
Anthropic
$5.0010.3M/mo$228Local cheaperDifferent model
Different-model comparison: the hosted model is not the model you would run locally. Treat this as a workload-quality trade-off, not a direct economic equivalence — the frontier model may complete a task in fewer tokens, or complete tasks the local model cannot. Break-even is the monthly output volume at which API spend equals the $134/month economic cost of owning this machine, assuming 8 input tokens per output token. Hosted prices are published in USD. This table is the canonical US default.