Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB)
A supported dual-GPU rack workstation that bridges the gap between a single 96 GB professional GPU and four-GPU deskside systems. Its 192 GB of aggregate ECC VRAM is sharded over PCIe rather than presented as one physically unified pool.
Dellsystem19 models fit0 measured
Memory
192 GB
Official spec GDDR7 ECC (distributed)
Memory bandwidth
3,584 GB/s
Official spec ~72% achieved in practice
Specifications
CPU
Dual Intel Xeon Scalable Official spec
GPU
2× 2x NVIDIA RTX PRO 6000 Blackwell Max-Q 96 GB Official spec
Memory
192 GB GDDR7 ECC (distributed) Official spec
Memory bandwidth
3,584 GB/s Official spec source
Achieved bandwidth ⓘ
~72% (2,580 GB/s) Assumption
FP16 compute ⓘ
~380 TFLOPS Official spec
System RAM
256 GB Official spec
Interconnect
PCIe 5.0 x16 Official spec
Architecture
multi_gpu Official spec
Nominal power
1200 W Official spec
Measured load power
920 W Measured
Idle power
135 W Measured
Released
1 Apr 2025 Official spec
Dell lists the dual 96 GB GPU option and a configurable rack platform; the Dutch VAT-inclusive system price is an ESTIMATE derived from the current US configurator. The GPUs shard models over PCIe, so 192 GB is aggregate capacity rather than a unified allocation. Power figures are engineering estimates, not measurements.
Standout models on this machine
Fastest
gpt-oss-20b — ~310–446 t/s
MXFP4 · estimated confidence
Largest that fits
Related hardware
Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB)384 GB · 7,168 GB/sCPU-only workstation (Ryzen 9950X, 192 GB DDR5)192 GB · 90 GB/sNVIDIA DGX H200 (8x H200, 1,128 GB)1,128 GB · 38,400 GB/sMac Studio M5 Ultra 256 GB256 GB · 1,200 GB/sQuad RTX 5090 workstation (4x 32 GB)128 GB · 7,168 GB/sMac Studio M3 Ultra 256 GB256 GB · 819 GB/s
Model performance on Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB)0 measured, 19 estimated
19 of 19 rows
| Model↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|
| gpt-oss-20b OpenAI · 20.915B (3.6B active) | MXFP4 | 13.4 GB | ~310–446 t/s | ~14800–30740 t/s | 128K | Comfortable | Estimated |
| gpt-oss-120b OpenAI · 116.829B (5.1B active) | MXFP4 | 66.9 GB | ~259–373 t/s | ~5260–10930 t/s | 128K | Comfortable | Estimated |
| Qwen3-Coder 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.9 GB | ~223–320 t/s | ~6400–13290 t/s | 256K | Comfortable | Estimated |
| Qwen3 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.9 GB | ~223–320 t/s | ~6400–13290 t/s | 40K | Comfortable | Estimated |
| Qwen3.8 Flash Next Alibaba Qwen · 180B (6B active) | Q4_K_M | 110.0 GB | ~218–313 t/s | ~1950–4060 t/s | 256K | Comfortable | Estimated |
| Gemma 4 26B-A4B Google DeepMind · 25.806B (3.8B active) | Q8_0 | 30.1 GB | ~203–293 t/s | ~6490–13470 t/s | 256K | Comfortable | Estimated |
| Qwen3 8B Alibaba Qwen · 8.191B | Q8_0 | 11.0 GB | ~132–190 t/s | ~7840–16280 t/s | 40K | Comfortable | Estimated |
| Qwen3.5 122B-A10B Alibaba Qwen · 125.086B (10B active) | Q8_0 | 132.2 GB | ~98.6–142 t/s | ~1820–3770 t/s | 192K | Comfortable | Estimated |
| Qwen3 235B-A22B Alibaba Qwen · 235.094B (22B active) | MLX 4-bit | 134.8 GB | ~86.7–125 t/s | ~1790–3710 t/s | 96K | Comfortable | Estimated |
| Phi-4 14B Microsoft · 14.66B | Q8_0 | 18.2 GB | ~81.8–118 t/s | ~4380–9100 t/s | 16K | Comfortable | Estimated |
| Devstral Small 24B Mistral AI · 23.572B | Q8_0 | 27.2 GB | ~53.7–77.3 t/s | ~2720–5660 t/s | 128K | Comfortable | Estimated |
| Mistral Small 3.2 24B Mistral AI · 24.011B | Q8_0 | 27.6 GB | ~52.8–76 t/s | ~2670–5550 t/s | 128K | Comfortable | Estimated |
| Gemma 3 27B Google DeepMind · 27.432B | Q8_0 | 33.8 GB | ~46.8–67.3 t/s | ~2340–4860 t/s | 128K | Comfortable | Estimated |
| Qwen3.8 27B Alibaba Qwen · 27.781B | Q8_0 | 32.3 GB | ~46.2–66.5 t/s | ~2310–4800 t/s | 256K | Comfortable | Estimated |
| Qwen3.6 27B Alibaba Qwen · 27.781B | Q8_0 | 32.3 GB | ~46.2–66.5 t/s | ~2310–4800 t/s | 256K | Comfortable | Estimated |
| Gemma 4 31B Google DeepMind · 31.273B | Q8_0 | 41.4 GB | ~41.4–59.6 t/s | ~2050–4270 t/s | 96K | Comfortable | Estimated |
| Qwen3 32B Alibaba Qwen · 32.762B | Q8_0 | 37.5 GB | ~39.7–57.1 t/s | ~1960–4070 t/s | 40K | Comfortable | Estimated |
| DeepSeek-R1-Distill 32B DeepSeek · 32.764B | Q8_0 | 37.5 GB | ~39.7–57.1 t/s | ~1960–4070 t/s | 128K | Comfortable | Estimated |
| Llama 3.3 70B Meta · 70.554B | Q8_0 | 77.2 GB | ~19.1–27.5 t/s | ~910–1890 t/s | 128K | Comfortable | Estimated |
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$467
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $367 more
Net cash at purchase
$36,663
Calculated VAT not reclaimable
Total over 5 years
$28,021
Calculated after tax, after resale
Cost per USD/1M tokens
$9.75
Calculated 47.9M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Monthly breakdown
| Depreciation 31,163 over 5 years, straight-line to a 5,499 residual | $519.39 |
| Electricity 51.4 kWh/month at 0.140/kWh | $7.19 |
| Cost of capital 4.0%/yr on 21,081 average capital employed | $70.27 |
| Monthly cost before tax | $596.85 |
| Electricity tax shield Running costs are deductible business expenses | −$1.51 |
| First-year expensing §179 (100.0%) | −$128.32 |
| Monthly economic cost after tax | $467.02 |
Three different numbers, deliberately
Cash cost
$36,663
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$519.39/month
$31,163 written down over 5 years to a $5,499 residual.
After-tax economic cost
$467.02/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$5,499 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
292 W Calculated
Load 920 W for 20% of powered hours, idle 135 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $7,699 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$70.27/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.