NVIDIA DGX H200 (8x H200, 1,128 GB)
An eight-GPU 8U server with 1,128 GB of aggregate HBM3e and NVSwitch. It is an enterprise ceiling for the catalogue: radically more expensive and power-hungry than a workstation, but able to keep very large open models entirely in accelerator memory.
NVIDIAappliance28 models fit0 measured
Memory
1,128 GB
Official spec HBM3e ECC (distributed)
Memory bandwidth
38,400 GB/s
Official spec ~78% achieved in practice
Specifications
CPU
2x Intel Xeon Platinum 8480C, 56-core Official spec
GPU
8× 8x NVIDIA H200 SXM 141 GB Official spec
Memory
1,128 GB HBM3e ECC (distributed) Official spec
Memory bandwidth
38,400 GB/s Official spec source
Achieved bandwidth ⓘ
~78% (29,952 GB/s) Assumption
FP16 compute ⓘ
~15,832 TFLOPS Official spec
System RAM
2048 GB Official spec
Interconnect
NVSwitch / 4th-generation NVLink, 900 GB/s per GPU Official spec
Architecture
multi_gpu Official spec
Nominal power
10200 W Official spec
Measured load power
8500 W Measured
Idle power
1600 W Measured
Released
1 Apr 2024 Official spec
Eight 141 GB H200 SXM GPUs provide 1,128 GB aggregate HBM3e and 38.4 TB/s aggregate bandwidth, connected through NVSwitch. Price and typical load are ESTIMATES; NVIDIA specifies an 8U system with a 10.2 kW maximum input. This belongs to the enterprise comparison tier rather than the prosumer shortlist.
Standout models on this machine
Fastest
gpt-oss-20b — ~553–796 t/s
MXFP4 · estimated confidence
Largest that fits
Related hardware
NVIDIA DGX Station GB300 (748 GB)748 GB · 7,100 GB/sLenovo ThinkStation PX (4x RTX PRO 6000, 384 GB)384 GB · 7,168 GB/sDell Precision 7960 Rack (2x RTX PRO 6000, 192 GB)192 GB · 3,584 GB/sQuad RTX 5090 workstation (4x 32 GB)128 GB · 7,168 GB/sNVIDIA DGX Spark 128 GB128 GB · 273 GB/sRTX PRO 6000 Max-Q workstation (96 GB, 300 W)96 GB · 1,792 GB/s
Model performance on NVIDIA DGX H200 (8x H200, 1,128 GB)0 measured, 28 estimated
28 of 28 rows
| Model↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|
| gpt-oss-20b OpenAI · 20.915B (3.6B active) | MXFP4 | 15.8 GB | ~553–796 t/s | ~616700–1280830 t/s | 128K | Comfortable | Estimated |
| gpt-oss-120b OpenAI · 116.829B (5.1B active) | MXFP4 | 69.3 GB | ~541–778 t/s | ~219230–455320 t/s | 128K | Comfortable | Estimated |
| Qwen3-Coder 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 36.3 GB | ~529–761 t/s | ~266560–553620 t/s | 256K | Comfortable | Estimated |
| Qwen3 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 36.3 GB | ~529–761 t/s | ~266560–553620 t/s | 40K | Comfortable | Estimated |
| Gemma 4 26B-A4B Google DeepMind · 25.806B (3.8B active) | Q8_0 | 32.5 GB | ~521–750 t/s | ~270190–561170 t/s | 256K | Comfortable | Estimated |
| Qwen3.8 Flash Next Alibaba Qwen · 180B (6B active) | Q8_0 | 191.7 GB | ~490–705 t/s | ~81420–169100 t/s | 256K | Comfortable | Estimated |
| Qwen3 8B Alibaba Qwen · 8.191B | Q8_0 | 13.4 GB | ~477–687 t/s | ~326650–678430 t/s | 40K | Comfortable | Estimated |
| DeepSeek-V4.1 Flash DeepSeek · 763.205B (16B active) | Q4_K_M | 461.0 GB | ~451–649 t/s | ~24210–50290 t/s | 1024K | Comfortable | Estimated |
| Qwen3.5 122B-A10B Alibaba Qwen · 125.086B (10B active) | Q8_0 | 134.6 GB | ~441–635 t/s | ~75650–157120 t/s | 256K | Comfortable | Estimated |
| Phi-4 14B Microsoft · 14.66B | Q8_0 | 20.6 GB | ~416–599 t/s | ~182510–379060 t/s | 16K | Comfortable | Estimated |
| DeepSeek-V4 Flash DeepSeek · 304.18B (13B active) | Q8_0 | 320.7 GB | ~411–592 t/s | ~42550–88370 t/s | 1024K | Comfortable | Estimated |
| GLM-5.3 Flash Z.ai · 321.323B (18B active) | Q8_0 | 338.2 GB | ~369–531 t/s | ~35180–73070 t/s | 1024K | Comfortable | Estimated |
| Kimi K2.6 Moonshot AI · 1.0T (32B active) | Q4_K_M | 619.4 GB | ~367–528 t/s | ~14760–30660 t/s | 256K | Comfortable | Estimated |
| Devstral Small 24B Mistral AI · 23.572B | Q8_0 | 29.6 GB | ~354–509 t/s | ~113510–235750 t/s | 128K | Comfortable | Estimated |
| Mistral Small 3.2 24B Mistral AI · 24.011B | Q8_0 | 30.0 GB | ~351–506 t/s | ~111430–231440 t/s | 128K | Comfortable | Estimated |
| Hunyuan Hy3 Tencent · 298.786B (21B active) | Q8_0 | 316.9 GB | ~347–500 t/s | ~33780–70150 t/s | 256K | Comfortable | Estimated |
| Qwen3 235B-A22B Alibaba Qwen · 235.094B (22B active) | Q8_0 | 249.7 GB | ~341–490 t/s | ~37200–77270 t/s | 256K | Comfortable | Estimated |
| MiniMax M3 MiniMax · 427.04B (23B active) | Q8_0 | 448.7 GB | ~334–481 t/s | ~27000–56070 t/s | 1024K | Comfortable | Estimated |
| Gemma 3 27B Google DeepMind · 27.432B | Q8_0 | 36.2 GB | ~332–478 t/s | ~97540–202570 t/s | 128K | Comfortable | Estimated |
| Qwen3.8 27B Alibaba Qwen · 27.781B | Q8_0 | 34.7 GB | ~331–476 t/s | ~96310–200030 t/s | 256K | Comfortable | Estimated |
| Qwen3.6 27B Alibaba Qwen · 27.781B | Q8_0 | 34.7 GB | ~331–476 t/s | ~96310–200030 t/s | 256K | Comfortable | Estimated |
| Gemma 4 31B Google DeepMind · 31.273B | Q8_0 | 43.8 GB | ~313–451 t/s | ~85560–177690 t/s | 256K | Comfortable | Estimated |
| Qwen3 32B Alibaba Qwen · 32.762B | Q8_0 | 39.9 GB | ~307–441 t/s | ~81670–169620 t/s | 40K | Comfortable | Estimated |
| DeepSeek-R1-Distill 32B DeepSeek · 32.764B | Q8_0 | 39.9 GB | ~307–441 t/s | ~81660–169610 t/s | 128K | Comfortable | Estimated |
| DeepSeek-V4 Pro DeepSeek · 1.6T (49B active) | Q4_K_M | 962.5 GB | ~306–441 t/s | ~9560–19850 t/s | — | Borderline | Estimated |
| DeepSeek-V3.2 DeepSeek · 685.397B (37B active) | Q8_0 | 716.9 GB | ~265–382 t/s | ~16800–34900 t/s | 160K | Comfortable | Estimated |
| GLM-5.3 Z.ai · 753.33B (40B active) | Q8_0 | 787.6 GB | ~254–365 t/s | ~15410–32010 t/s | 1024K | Comfortable | Estimated |
| Llama 3.3 70B Meta · 70.554B | Q8_0 | 79.6 GB | ~198–285 t/s | ~37920–78760 t/s | 128K | Comfortable | Estimated |
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$4,743
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $4,643 more
Net cash at purchase
$359,430
Calculated VAT not reclaimable
Total over 5 years
$284,555
Calculated after tax, after resale
Cost per USD/1M tokens
$55.48
Calculated 85.5M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Monthly breakdown
| Depreciation 316,298 over 5 years, straight-line to a 43,132 residual | $5,271.64 |
| Electricity 524.5 kWh/month at 0.140/kWh | $73.43 |
| Cost of capital 4.0%/yr on 201,281 average capital employed | $670.94 |
| Monthly cost before tax | $6,016.00 |
| Electricity tax shield Running costs are deductible business expenses | −$15.42 |
| First-year expensing §179 (100.0%) | −$1,258.00 |
| Monthly economic cost after tax | $4,742.58 |
Three different numbers, deliberately
Cash cost
$359,430
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$5,271.64/month
$316,298 written down over 5 years to a $43,132 residual.
After-tax economic cost
$4,742.58/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$43,132 (12%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
2980 W Calculated
Load 8500 W for 20% of powered hours, idle 1600 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $75,480 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$670.94/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.