RTX PRO 6000 Max-Q workstation (96 GB, 300 W)

96 GB of GDDR7 at 1,792 GB/s in a single card. The only way to get both dense-model bandwidth and enough memory for a 100B-class model in one PCIe slot, at a price that reflects exactly that.

NVIDIAcomponent17 models fit0 measured
Memory
96 GB
Official spec GDDR7 ECC
Memory bandwidth
1,792 GB/s
Official spec ~74% achieved in practice
Specifications
CPU
AMD Ryzen 9 9950X Official spec
GPU
RTX PRO 6000 Blackwell Max-Q 96 GB Official spec
Memory
96 GB GDDR7 ECC Official spec
Memory bandwidth
1,792 GB/s Official spec source
Achieved bandwidth
~74% (1,332 GB/s) Assumption
FP16 compute
~190 TFLOPS Official spec
System RAM
128 GB Official spec
Architecture
Discrete GPU with dedicated VRAM Official spec
Nominal power
450 W Official spec
Measured load power
390 W Measured
Idle power
45 W Measured
Released
1 Apr 2025 Official spec
Same 96 GB and same memory bus at a 300 W card power limit. Decode barely changes because decode is bandwidth-bound, not compute-bound — this is the clearest illustration in the database of why the two metrics must be modelled separately. Achieved-bandwidth coefficient INHERITED from rtx-pro-6000-workstation, which shares the same silicon and software stack.
Model performance on RTX PRO 6000 Max-Q workstation (96 GB, 300 W)0 measured, 17 estimated
17 of 17 rows
ModelQuantMemoryDecodePrefillContextFitConfidence
gpt-oss-20b
OpenAI · 20.915B (3.6B active)
MXFP413.0 GB~329–474 t/s~6140–12760 t/s128KComfortableEstimated
gpt-oss-120b
OpenAI · 116.829B (5.1B active)
MXFP466.5 GB~257–369 t/s~2180–4530 t/s128KComfortableEstimated
Qwen3-Coder 30B-A3B
Alibaba Qwen · 30.532B (3.3B active)
Q8_033.5 GB~210–303 t/s~2650–5510 t/s256KComfortableEstimated
Qwen3 30B-A3B
Alibaba Qwen · 30.532B (3.3B active)
Q8_033.5 GB~210–303 t/s~2650–5510 t/s40KComfortableEstimated
Gemma 4 26B-A4B
Google DeepMind · 25.806B (3.8B active)
Q8_029.7 GB~188–270 t/s~2690–5590 t/s192KComfortableEstimated
Qwen3.5 122B-A10B
Alibaba Qwen · 125.086B (10B active)
Q4_K_M76.7 GB~133–192 t/s~753–1560 t/s32KFitsEstimated
Qwen3 8B
Alibaba Qwen · 8.191B
Q8_010.6 GB~112–161 t/s~3250–6760 t/s40KComfortableEstimated
Phi-4 14B
Microsoft · 14.66B
Q8_017.8 GB~65.6–94.4 t/s~1820–3780 t/s16KComfortableEstimated
Devstral Small 24B
Mistral AI · 23.572B
Q8_026.8 GB~41.8–60.2 t/s~1130–2350 t/s128KComfortableEstimated
Mistral Small 3.2 24B
Mistral AI · 24.011B
Q8_027.2 GB~41.1–59.1 t/s~1110–2310 t/s128KComfortableEstimated
Gemma 3 27B
Google DeepMind · 27.432B
Q8_033.4 GB~36.1–52 t/s~971–2020 t/s96KComfortableEstimated
Qwen3.8 27B
Alibaba Qwen · 27.781B
Q8_031.9 GB~35.7–51.4 t/s~959–1990 t/s192KComfortableEstimated
Qwen3.6 27B
Alibaba Qwen · 27.781B
Q8_031.9 GB~35.7–51.4 t/s~959–1990 t/s192KComfortableEstimated
Gemma 4 31B
Google DeepMind · 31.273B
Q8_041.0 GB~31.8–45.8 t/s~852–1770 t/s32KComfortableEstimated
Qwen3 32B
Alibaba Qwen · 32.762B
Q8_037.1 GB~30.4–43.8 t/s~813–1690 t/s40KComfortableEstimated
DeepSeek-R1-Distill 32B
DeepSeek · 32.764B
Q8_037.1 GB~30.4–43.8 t/s~813–1690 t/s128KComfortableEstimated
Llama 3.3 70B
Meta · 70.554B
Q4_K_M45.8 GB~24.9–35.8 t/s~378–784 t/s96KComfortableEstimated
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$138
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $38 more
Net cash at purchase
$10,809
Calculated VAT not reclaimable
Total over 5 years
$8,294
Calculated after tax, after resale
Cost per USD/1M tokens
$2.72
Calculated 50.9M tokens/month

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Monthly breakdown

Depreciation
9,188 over 5 years, straight-line to a 1,621 residual
$153.13
Electricity
20.1 kWh/month at 0.140/kWh
$2.81
Cost of capital
4.0%/yr on 6,215 average capital employed
$20.72
Monthly cost before tax$176.65
Electricity tax shield
Running costs are deductible business expenses
−$0.59
First-year expensing
§179 (100.0%)
−$37.83
Monthly economic cost after tax$138.23

Three different numbers, deliberately

Cash cost
$10,809
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$153.13/month
$9,188 written down over 5 years to a $1,621 residual.
After-tax economic cost
$138.23/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$1,621 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
114 W Calculated
Load 390 W for 20% of powered hours, idle 45 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $2,270 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$20.72/month Assumption
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.