gpt-oss-120b on NVIDIA DGX Spark 128 GB

116.829B (5.1B active) on 128 GB at 273 GB/s. Hardware details · Model details

CompatibilityComfortable
Fits
Yes
Calculated Uses under 80% of usable memory. Room for a long context and other work.
Recommended quantization
MXFP4
Calculated gguf
Memory required
66.5 GB
Calculated of 105.0 GB usable — 63%
Max practical context
128K
Calculated model supports 128K

Memory budget at 8K context

Model weights
63.0 GB Calculated
KV cache
0.6 GB Calculated
Runtime overhead
2.9 GB Estimated
Total required
66.5 GB Calculated
Usable memory
105.0 GB Assumption
Headroom
38.5 GB Calculated
Highest-precision quantization that leaves headroom: uses 63% of usable memory at 8K context.
Unified LPDDR: the accelerator allocates from the same 128 GB pool as the OS. We assume 82% is available to the inference engine.
Mixture of experts: all 116.829B parameters must be resident in memory even though only ~5.1B are active per token. Memory follows total parameters; speed follows active parameters.
PerformanceHigh9/10
Decode (generation)
56.4 t/s
Aggregated range 51.5–60.6 across 4 runs
Prefill (prompt)
1680 t/s
Aggregated range 1,512–1,956
TTFT at 8K
~3.9–8.2 s
Estimated time to first token
Power while generating
170 W
Measured 1.04 tokens/s per watt
Backed by 4 published measurements. Every run is listed below with its source, engine version and context depth.

Evidence

TypeEngineContextDecodePrefillDateSource
Measured (community)llama.cpp b7db35a7060.57 t/s1,956 t/s20 Oct 2025github
Measured (independent)llama.cpp058.70 t/s1,723 t/s25 Oct 2025blog
Measured (community)llama.cpp b7db35a74K54.14 t/s1,637 t/s20 Oct 2025github
Measured (community)llama.cpp b7db35a78K51.54 t/s1,512 t/s20 Oct 2025github
Measured (community)llama.cpp b7db35a716K47.45 t/s1,307 t/s20 Oct 2025github
Measured (community)llama.cpp b7db35a732K40.55 t/s1,027 t/s20 Oct 2025github
  • llama-bench at zero context depth: the best case, and the number most often quoted in marketing. Reported as 60.57 +/- 0.25.
  • Same run at 4K context depth. Reported as 54.14 +/- 0.08.
  • 8K context depth. Reported as 51.54 +/- 0.14.
  • 16K context depth. Reported as 47.45 +/- 0.08.
  • 32K context depth — decode has fallen 33% from the zero-context figure. This is why the site buckets by context rather than quoting one number.
  • Independent reproduction. Slightly below the llama.cpp thread figure, which is normal for a different kernel build.
How the estimate is calculated
  1. Decode: reading 5.1B active parameters at 4.25 bits/weight takes 14.81 ms at 273 GB/s x 67% achieved efficiency.
  2. MoE routing penalty of 15% applied: expert gathers are less bandwidth-efficient than a dense sweep.
  3. Prefill: 125 TFLOPS (FP16) x 4 for native FP4 tensor cores x 0.85 calibrated against measured prefill on this platform x 18% assumed model-FLOPs utilisation, divided by 2 x 24.4B parameters per token.
  4. Prefill uses 24.4B effective parameters, not the 5.1B active in decode: a batch of hundreds of tokens routes across most of the expert pool.

  • MoE decode depends on how well the engine batches expert gathers; real results vary more than for dense models.
  • Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
  • This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.

Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.

Why this is rated "High" confidence
  • Independent third-party benchmark: +3
  • 4 measurements: +2
  • Sources agree closely (±7%): +2
  • Complete run metadata (engine, version, context, source, date): +2

Score 9/10. Aggregated with the median, not the mean, so one outlier cannot move the headline figure.

This model at other quantizations on this machine
QuantizationWeightsTotal neededUtilisationFitMax context
MXFP4recommended63.0 GB66.5 GB63%Comfortable128K
BF16217.6 GB225.7 GB215%Does not fit0
Cost of running gpt-oss-120b on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$52
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $48 less
Net cash at purchase
$3,873
Calculated VAT not reclaimable
Total over 5 years
$3,091
Calculated after tax, after resale
Cost per USD/1M tokens
$7.21
Calculated 7.1M tokens/month

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Monthly breakdown

Depreciation
3,408 over 5 years, straight-line to a 465 residual
$56.80
Electricity
9.5 kWh/month at 0.140/kWh
$1.33
Cost of capital
4.0%/yr on 2,169 average capital employed
$7.23
Monthly cost before tax$65.36
Electricity tax shield
Running costs are deductible business expenses
−$0.28
First-year expensing
§179 (100.0%)
−$13.55
Monthly economic cost after tax$51.52

Three different numbers, deliberately

Cash cost
$3,873
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$56.80/month
$3,408 written down over 5 years to a $465 residual.
After-tax economic cost
$51.52/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$465 (12%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
54 W Calculated
Load 170 W for 20% of powered hours, idle 25 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $813 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$7.23/month Assumption
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.
Efficiency metrics
Decode per $1,000 spent
14.6 t/s
Calculated purchase price only
Decode per $100/month
109.5 t/s
Calculated after-tax ownership cost
Tokens per joule
1.04
Calculated same as tokens/s per watt
USD per 1M output tokens
$7.21
Calculated at 20% utilisation
Local versus hosted APIs
Hosted modelUSD/1M outputBreak-evenAPI at your volumeVerdictComparison type
Qwen3.8 Flash
Alibaba Cloud
$0.4233.5M/mo$11API cheaperDifferent model
DeepSeek-V4 Flash
DeepSeek
$0.6621.3M/mo$17API cheaperDifferent model
DeepSeek-V4 Pro
DeepSeek
$1.987.1M/mo$52About equalDifferent model
DeepSeek V4 Pro (Together)
Together AI
$4.402.4M/mo$152Local cheaperDifferent model
GLM-5.3
Z.ai
$4.403.3M/mo$112Local cheaperDifferent model
Claude Haiku 4.5
Anthropic
$5.004.0M/mo$93Local cheaperDifferent model
Different-model comparison: the hosted model is not the model you would run locally. Treat this as a workload-quality trade-off, not a direct economic equivalence — the frontier model may complete a task in fewer tokens, or complete tasks the local model cannot. Break-even is the monthly output volume at which API spend equals the $52/month economic cost of owning this machine, assuming 8 input tokens per output token. Hosted prices are published in USD. This table is the canonical US default.
Page quality score (why this page is or is not indexed)

12/13 indexed. Generated pages are gated so we do not ask a search engine to rank a page with nothing computed to say. The directive is emitted in the page head via the metadata API, not in the body, so it is authoritative.

  • Has a memory-fit calculation (+3)
  • Has at least one real measurement (+4)
  • Has a price (+2)
  • Has an ownership economics calculation (+2)
  • Well connected (16 internal links) (+1)