RTX PRO 6000 Blackwell workstation (96 GB)
96 GB of GDDR7 at 1,792 GB/s in a single card. The only way to get both dense-model bandwidth and enough memory for a 100B-class model in one PCIe slot, at a price that reflects exactly that.
NVIDIAcomponent17 models fit0 measured
Memory
96 GB
Official spec GDDR7 ECC
Memory bandwidth
1,792 GB/s
Official spec ~74% achieved in practice
Specifications
CPU
AMD Ryzen 9 9950X Official spec
GPU
RTX PRO 6000 Blackwell 96 GB Official spec
Memory
96 GB GDDR7 ECC Official spec
Memory bandwidth
1,792 GB/s Official spec source
Achieved bandwidth ⓘ
~74% (1,332 GB/s) Assumption
FP16 compute ⓘ
~232 TFLOPS Official spec
System RAM
128 GB Official spec
Architecture
Discrete GPU with dedicated VRAM Official spec
Nominal power
750 W Official spec
Measured load power
660 W Measured
Idle power
60 W Measured
Released
1 Apr 2025 Official spec
The only single card that combines Mac-Studio-class capacity with RTX-class bandwidth. On decode it is roughly 2.5x an M5 Ultra for models that fit in 96 GB, at roughly 3x the power draw. Achieved-bandwidth coefficient CALIBRATED to 74% from 1 measured run on this configuration.
Standout models on this machine
Fastest
gpt-oss-20b — ~329–474 t/s
MXFP4 · estimated confidence
Largest that fits
Compare against
Related hardware
RTX PRO 6000 Max-Q workstation (96 GB, 300 W)96 GB · 1,792 GB/sQuad RTX 5090 workstation (4x 32 GB)128 GB · 7,168 GB/sDual RTX 5090 workstation (2x 32 GB)64 GB · 3,584 GB/sDual used RTX 4090 workstation (48 GB)48 GB · 2,016 GB/sRTX 5090 workstation (1x 32 GB)32 GB · 1,792 GB/sUsed RTX 4090 workstation (24 GB)24 GB · 1,008 GB/s
Model performance on RTX PRO 6000 Blackwell workstation (96 GB)0 measured, 17 estimated
17 of 17 rows
| Model↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|
| gpt-oss-20b OpenAI · 20.915B (3.6B active) | MXFP4 | 13.0 GB | ~329–474 t/s | ~7500–15580 t/s | 128K | Comfortable | Estimated |
| gpt-oss-120b OpenAI · 116.829B (5.1B active) | MXFP4 | 66.5 GB | ~257–369 t/s | ~2670–5540 t/s | 128K | Comfortable | Estimated |
| Qwen3-Coder 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~210–303 t/s | ~3240–6730 t/s | 256K | Comfortable | Estimated |
| Qwen3 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~210–303 t/s | ~3240–6730 t/s | 40K | Comfortable | Estimated |
| Gemma 4 26B-A4B Google DeepMind · 25.806B (3.8B active) | Q8_0 | 29.7 GB | ~188–270 t/s | ~3290–6820 t/s | 192K | Comfortable | Estimated |
| Qwen3.5 122B-A10B Alibaba Qwen · 125.086B (10B active) | Q4_K_M | 76.7 GB | ~133–192 t/s | ~920–1910 t/s | 32K | Fits | Estimated |
| Qwen3 8B Alibaba Qwen · 8.191B | Q8_0 | 10.6 GB | ~112–161 t/s | ~3970–8250 t/s | 40K | Comfortable | Estimated |
| Phi-4 14B Microsoft · 14.66B | Q8_0 | 17.8 GB | ~65.6–94.4 t/s | ~2220–4610 t/s | 16K | Comfortable | Estimated |
| Devstral Small 24B Mistral AI · 23.572B | Q8_0 | 26.8 GB | ~41.8–60.2 t/s | ~1380–2870 t/s | 128K | Comfortable | Estimated |
| Mistral Small 3.2 24B Mistral AI · 24.011B | Q8_0 | 27.2 GB | ~41.1–59.1 t/s | ~1360–2810 t/s | 128K | Comfortable | Estimated |
| Gemma 3 27B Google DeepMind · 27.432B | Q8_0 | 33.4 GB | ~36.1–52 t/s | ~1190–2460 t/s | 96K | Comfortable | Estimated |
| Qwen3.8 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~35.7–51.4 t/s | ~1170–2430 t/s | 192K | Comfortable | Estimated |
| Qwen3.6 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~35.7–51.4 t/s | ~1170–2430 t/s | 192K | Comfortable | Estimated |
| Gemma 4 31B Google DeepMind · 31.273B | Q8_0 | 41.0 GB | ~31.8–45.8 t/s | ~1040–2160 t/s | 32K | Comfortable | Estimated |
| Qwen3 32B Alibaba Qwen · 32.762B | Q8_0 | 37.1 GB | ~30.4–43.8 t/s | ~993–2060 t/s | 40K | Comfortable | Estimated |
| DeepSeek-R1-Distill 32B DeepSeek · 32.764B | Q8_0 | 37.1 GB | ~30.4–43.8 t/s | ~993–2060 t/s | 128K | Comfortable | Estimated |
| Llama 3.3 70B Meta · 70.554B | Q4_K_M | 45.8 GB | ~24.9–35.8 t/s | ~461–958 t/s | 96K | Comfortable | Estimated |
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$134
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $34 more
Net cash at purchase
$10,359
Calculated VAT not reclaimable
Total over 5 years
$8,031
Calculated after tax, after resale
Cost per USD/1M tokens
$2.63
Calculated 50.9M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Monthly breakdown
| Depreciation 8,805 over 5 years, straight-line to a 1,554 residual | $146.75 |
| Electricity 31.7 kWh/month at 0.140/kWh | $4.44 |
| Cost of capital 4.0%/yr on 5,956 average capital employed | $19.85 |
| Monthly cost before tax | $171.04 |
| Electricity tax shield Running costs are deductible business expenses | −$0.93 |
| First-year expensing §179 (100.0%) | −$36.26 |
| Monthly economic cost after tax | $133.85 |
Three different numbers, deliberately
Cash cost
$10,359
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$146.75/month
$8,805 written down over 5 years to a $1,554 residual.
After-tax economic cost
$133.85/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$1,554 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
180 W Calculated
Load 660 W for 20% of powered hours, idle 60 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $2,175 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$19.85/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.