NVIDIA DGX H200 (8x H200, 1,128 GB)

An eight-GPU 8U server with 1,128 GB of aggregate HBM3e and NVSwitch. It is an enterprise ceiling for the catalogue: radically more expensive and power-hungry than a workstation, but able to keep very large open models entirely in accelerator memory.

NVIDIAappliance28 models fit0 measured
Memory
1,128 GB
Official spec HBM3e ECC (distributed)
Memory bandwidth
38,400 GB/s
Official spec ~78% achieved in practice
Specifications
CPU
2x Intel Xeon Platinum 8480C, 56-core Official spec
GPU
8× 8x NVIDIA H200 SXM 141 GB Official spec
Memory
1,128 GB HBM3e ECC (distributed) Official spec
Memory bandwidth
38,400 GB/s Official spec source
Achieved bandwidth
~78% (29,952 GB/s) Assumption
FP16 compute
~15,832 TFLOPS Official spec
System RAM
2048 GB Official spec
Interconnect
NVSwitch / 4th-generation NVLink, 900 GB/s per GPU Official spec
Architecture
multi_gpu Official spec
Nominal power
10200 W Official spec
Measured load power
8500 W Measured
Idle power
1600 W Measured
Released
1 Apr 2024 Official spec
Eight 141 GB H200 SXM GPUs provide 1,128 GB aggregate HBM3e and 38.4 TB/s aggregate bandwidth, connected through NVSwitch. Price and typical load are ESTIMATES; NVIDIA specifies an 8U system with a 10.2 kW maximum input. This belongs to the enterprise comparison tier rather than the prosumer shortlist.
Model performance on NVIDIA DGX H200 (8x H200, 1,128 GB)0 measured, 28 estimated
28 of 28 rows
ModelQuantMemoryDecodePrefillContextFitConfidence
gpt-oss-20b
OpenAI · 20.915B (3.6B active)
MXFP415.8 GB~553–796 t/s~616700–1280830 t/s128KComfortableEstimated
gpt-oss-120b
OpenAI · 116.829B (5.1B active)
MXFP469.3 GB~541–778 t/s~219230–455320 t/s128KComfortableEstimated
Qwen3-Coder 30B-A3B
Alibaba Qwen · 30.532B (3.3B active)
Q8_036.3 GB~529–761 t/s~266560–553620 t/s256KComfortableEstimated
Qwen3 30B-A3B
Alibaba Qwen · 30.532B (3.3B active)
Q8_036.3 GB~529–761 t/s~266560–553620 t/s40KComfortableEstimated
Gemma 4 26B-A4B
Google DeepMind · 25.806B (3.8B active)
Q8_032.5 GB~521–750 t/s~270190–561170 t/s256KComfortableEstimated
Qwen3.8 Flash Next
Alibaba Qwen · 180B (6B active)
Q8_0191.7 GB~490–705 t/s~81420–169100 t/s256KComfortableEstimated
Qwen3 8B
Alibaba Qwen · 8.191B
Q8_013.4 GB~477–687 t/s~326650–678430 t/s40KComfortableEstimated
DeepSeek-V4.1 Flash
DeepSeek · 763.205B (16B active)
Q4_K_M461.0 GB~451–649 t/s~24210–50290 t/s1024KComfortableEstimated
Qwen3.5 122B-A10B
Alibaba Qwen · 125.086B (10B active)
Q8_0134.6 GB~441–635 t/s~75650–157120 t/s256KComfortableEstimated
Phi-4 14B
Microsoft · 14.66B
Q8_020.6 GB~416–599 t/s~182510–379060 t/s16KComfortableEstimated
DeepSeek-V4 Flash
DeepSeek · 304.18B (13B active)
Q8_0320.7 GB~411–592 t/s~42550–88370 t/s1024KComfortableEstimated
GLM-5.3 Flash
Z.ai · 321.323B (18B active)
Q8_0338.2 GB~369–531 t/s~35180–73070 t/s1024KComfortableEstimated
Kimi K2.6
Moonshot AI · 1.0T (32B active)
Q4_K_M619.4 GB~367–528 t/s~14760–30660 t/s256KComfortableEstimated
Devstral Small 24B
Mistral AI · 23.572B
Q8_029.6 GB~354–509 t/s~113510–235750 t/s128KComfortableEstimated
Mistral Small 3.2 24B
Mistral AI · 24.011B
Q8_030.0 GB~351–506 t/s~111430–231440 t/s128KComfortableEstimated
Hunyuan Hy3
Tencent · 298.786B (21B active)
Q8_0316.9 GB~347–500 t/s~33780–70150 t/s256KComfortableEstimated
Qwen3 235B-A22B
Alibaba Qwen · 235.094B (22B active)
Q8_0249.7 GB~341–490 t/s~37200–77270 t/s256KComfortableEstimated
MiniMax M3
MiniMax · 427.04B (23B active)
Q8_0448.7 GB~334–481 t/s~27000–56070 t/s1024KComfortableEstimated
Gemma 3 27B
Google DeepMind · 27.432B
Q8_036.2 GB~332–478 t/s~97540–202570 t/s128KComfortableEstimated
Qwen3.8 27B
Alibaba Qwen · 27.781B
Q8_034.7 GB~331–476 t/s~96310–200030 t/s256KComfortableEstimated
Qwen3.6 27B
Alibaba Qwen · 27.781B
Q8_034.7 GB~331–476 t/s~96310–200030 t/s256KComfortableEstimated
Gemma 4 31B
Google DeepMind · 31.273B
Q8_043.8 GB~313–451 t/s~85560–177690 t/s256KComfortableEstimated
Qwen3 32B
Alibaba Qwen · 32.762B
Q8_039.9 GB~307–441 t/s~81670–169620 t/s40KComfortableEstimated
DeepSeek-R1-Distill 32B
DeepSeek · 32.764B
Q8_039.9 GB~307–441 t/s~81660–169610 t/s128KComfortableEstimated
DeepSeek-V4 Pro
DeepSeek · 1.6T (49B active)
Q4_K_M962.5 GB~306–441 t/s~9560–19850 t/sBorderlineEstimated
DeepSeek-V3.2
DeepSeek · 685.397B (37B active)
Q8_0716.9 GB~265–382 t/s~16800–34900 t/s160KComfortableEstimated
GLM-5.3
Z.ai · 753.33B (40B active)
Q8_0787.6 GB~254–365 t/s~15410–32010 t/s1024KComfortableEstimated
Llama 3.3 70B
Meta · 70.554B
Q8_079.6 GB~198–285 t/s~37920–78760 t/s128KComfortableEstimated
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$4,743
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $4,643 more
Net cash at purchase
$359,430
Calculated VAT not reclaimable
Total over 5 years
$284,555
Calculated after tax, after resale
Cost per USD/1M tokens
$55.48
Calculated 85.5M tokens/month

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Monthly breakdown

Depreciation
316,298 over 5 years, straight-line to a 43,132 residual
$5,271.64
Electricity
524.5 kWh/month at 0.140/kWh
$73.43
Cost of capital
4.0%/yr on 201,281 average capital employed
$670.94
Monthly cost before tax$6,016.00
Electricity tax shield
Running costs are deductible business expenses
−$15.42
First-year expensing
§179 (100.0%)
−$1,258.00
Monthly economic cost after tax$4,742.58

Three different numbers, deliberately

Cash cost
$359,430
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$5,271.64/month
$316,298 written down over 5 years to a $43,132 residual.
After-tax economic cost
$4,742.58/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$43,132 (12%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
2980 W Calculated
Load 8500 W for 20% of powered hours, idle 1600 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $75,480 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$670.94/month Assumption
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.