NVIDIA DGX Station GB300 (748 GB)
A deskside Grace Blackwell Ultra system with 252 GB of HBM3e and 496 GB of LPDDR5X joined into 748 GB of coherent memory over NVLink-C2C. It is the first desktop-sized NVIDIA system in the catalogue that can hold trillion-parameter-class models at aggressive quantization.
NVIDIAappliance27 models fit0 measured
Memory
748 GB
Official spec 252 GB HBM3e + 496 GB LPDDR5X coherent
Memory bandwidth
7,100 GB/s
Official spec ~15% achieved in practice
Specifications
CPU
NVIDIA Grace, 72-core Arm Official spec
GPU
NVIDIA GB300 Blackwell Ultra Official spec
Memory
748 GB 252 GB HBM3e + 496 GB LPDDR5X coherent Official spec
Memory bandwidth
7,100 GB/s Official spec source
Achieved bandwidth ⓘ
~15% (1,065 GB/s) Assumption
FP16 compute ⓘ
~2,500 TFLOPS Official spec
Interconnect
NVLink-C2C coherent memory Official spec
Architecture
Unified memory (CPU and GPU share one pool) Official spec
Nominal power
1600 W Official spec
Measured load power
1350 W Measured
Idle power
180 W Measured
Released
28 Aug 2026 Official spec
Official capacity is 748 GB coherent: 252 GB HBM3e at 7.1 TB/s plus 496 GB CPU LPDDR5X over NVLink-C2C. The displayed 7.1 TB/s is the fast HBM tier, not uniform bandwidth across the full pool; the deliberately conservative achieved-bandwidth coefficient accounts for models spilling into LPDDR. Price and typical draw are ESTIMATES; NVIDIA specifies a 1,600 W system power budget.
Standout models on this machine
Fastest
gpt-oss-20b — ~263–379 t/s
MXFP4 · estimated confidence
Largest that fits
Related hardware
Mac Studio M5 Ultra 512 GB512 GB · 1,200 GB/sMac Studio M3 Ultra 512 GB512 GB · 819 GB/sNVIDIA DGX H200 (8x H200, 1,128 GB)1,128 GB · 38,400 GB/sMac Studio M5 Ultra 256 GB256 GB · 1,200 GB/sMac Studio M3 Ultra 256 GB256 GB · 819 GB/sLenovo ThinkStation PX (4x RTX PRO 6000, 384 GB)384 GB · 7,168 GB/s
Model performance on NVIDIA DGX Station GB300 (748 GB)0 measured, 27 estimated
27 of 27 rows
| Model↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|
| gpt-oss-20b OpenAI · 20.915B (3.6B active) | MXFP4 | 13.0 GB | ~263–379 t/s | ~82400–171140 t/s | 128K | Comfortable | Estimated |
| gpt-oss-120b OpenAI · 116.829B (5.1B active) | MXFP4 | 66.5 GB | ~205–296 t/s | ~29290–60840 t/s | 128K | Comfortable | Estimated |
| Qwen3-Coder 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~168–242 t/s | ~35620–73970 t/s | 256K | Comfortable | Estimated |
| Qwen3 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~168–242 t/s | ~35620–73970 t/s | 40K | Comfortable | Estimated |
| Gemma 4 26B-A4B Google DeepMind · 25.806B (3.8B active) | Q8_0 | 29.7 GB | ~150–216 t/s | ~36100–74980 t/s | 256K | Comfortable | Estimated |
| Qwen3.8 Flash Next Alibaba Qwen · 180B (6B active) | Q8_0 | 188.9 GB | ~102–147 t/s | ~10880–22590 t/s | 256K | Comfortable | Estimated |
| Qwen3 8B Alibaba Qwen · 8.191B | Q8_0 | 10.6 GB | ~89.4–129 t/s | ~43650–90650 t/s | 40K | Comfortable | Estimated |
| DeepSeek-V4.1 Flash DeepSeek · 763.205B (16B active) | Q4_K_M | 458.2 GB | ~70–101 t/s | ~3240–6720 t/s | 1024K | Comfortable | Estimated |
| Qwen3.5 122B-A10B Alibaba Qwen · 125.086B (10B active) | Q8_0 | 131.8 GB | ~64.4–92.6 t/s | ~10110–20990 t/s | 256K | Comfortable | Estimated |
| Phi-4 14B Microsoft · 14.66B | Q8_0 | 17.8 GB | ~52.5–75.5 t/s | ~24390–50650 t/s | 16K | Comfortable | Estimated |
| DeepSeek-V4 Flash DeepSeek · 304.18B (13B active) | Q8_0 | 317.9 GB | ~50.4–72.6 t/s | ~5690–11810 t/s | 1024K | Comfortable | Estimated |
| GLM-5.3 Flash Z.ai · 321.323B (18B active) | Q8_0 | 335.4 GB | ~37.1–53.3 t/s | ~4700–9760 t/s | 1024K | Comfortable | Estimated |
| Kimi K2.6 Moonshot AI · 1.0T (32B active) | Q4_K_M | 616.6 GB | ~36.6–52.6 t/s | ~1970–4100 t/s | — | Borderline | Estimated |
| Devstral Small 24B Mistral AI · 23.572B | Q8_0 | 26.8 GB | ~33.4–48.1 t/s | ~15170–31500 t/s | 128K | Comfortable | Estimated |
| Mistral Small 3.2 24B Mistral AI · 24.011B | Q8_0 | 27.2 GB | ~32.9–47.3 t/s | ~14890–30920 t/s | 128K | Comfortable | Estimated |
| Hunyuan Hy3 Tencent · 298.786B (21B active) | Q8_0 | 314.1 GB | ~32–46 t/s | ~4510–9370 t/s | 256K | Comfortable | Estimated |
| DeepSeek-V3.2 DeepSeek · 685.397B (37B active) | Q4_K_M | 412.1 GB | ~31.8–45.8 t/s | ~2240–4660 t/s | 160K | Comfortable | Estimated |
| Qwen3 235B-A22B Alibaba Qwen · 235.094B (22B active) | Q8_0 | 246.9 GB | ~30.6–44 t/s | ~4970–10320 t/s | 256K | Comfortable | Estimated |
| GLM-5.3 Z.ai · 753.33B (40B active) | Q4_K_M | 452.9 GB | ~29.5–42.5 t/s | ~2060–4280 t/s | 1024K | Comfortable | Estimated |
| MiniMax M3 MiniMax · 427.04B (23B active) | Q8_0 | 445.9 GB | ~29.3–42.1 t/s | ~3610–7490 t/s | 1024K | Comfortable | Estimated |
| Gemma 3 27B Google DeepMind · 27.432B | Q8_0 | 33.4 GB | ~28.9–41.6 t/s | ~13030–27070 t/s | 128K | Comfortable | Estimated |
| Qwen3.8 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~28.6–41.1 t/s | ~12870–26730 t/s | 256K | Comfortable | Estimated |
| Qwen3.6 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~28.6–41.1 t/s | ~12870–26730 t/s | 256K | Comfortable | Estimated |
| Gemma 4 31B Google DeepMind · 31.273B | Q8_0 | 41.0 GB | ~25.5–36.6 t/s | ~11430–23740 t/s | 256K | Comfortable | Estimated |
| Qwen3 32B Alibaba Qwen · 32.762B | Q8_0 | 37.1 GB | ~24.3–35 t/s | ~10910–22660 t/s | 40K | Comfortable | Estimated |
| DeepSeek-R1-Distill 32B DeepSeek · 32.764B | Q8_0 | 37.1 GB | ~24.3–35 t/s | ~10910–22660 t/s | 128K | Comfortable | Estimated |
| Llama 3.3 70B Meta · 70.554B | Q8_0 | 76.8 GB | ~11.5–16.5 t/s | ~5070–10520 t/s | 128K | Comfortable | Estimated |
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$1,182
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $1,082 more
Net cash at purchase
$90,082
Calculated VAT not reclaimable
Total over 5 years
$70,927
Calculated after tax, after resale
Cost per USD/1M tokens
$29.04
Calculated 40.7M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Monthly breakdown
| Depreciation 79,272 over 5 years, straight-line to a 10,810 residual | $1,321.20 |
| Electricity 72.9 kWh/month at 0.140/kWh | $10.20 |
| Cost of capital 4.0%/yr on 50,446 average capital employed | $168.15 |
| Monthly cost before tax | $1,499.55 |
| Electricity tax shield Running costs are deductible business expenses | −$2.14 |
| First-year expensing §179 (100.0%) | −$315.29 |
| Monthly economic cost after tax | $1,182.12 |
Three different numbers, deliberately
Cash cost
$90,082
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$1,321.20/month
$79,272 written down over 5 years to a $10,810 residual.
After-tax economic cost
$1,182.12/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$10,810 (12%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
414 W Calculated
Load 1350 W for 20% of powered hours, idle 180 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $18,917 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$168.15/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.