Dual RTX 5090 workstation (2x 32 GB) vs Mac Studio M5 Ultra 256 GB

The Dual RTX 5090 is about 57% faster across 7 like-for-like model configurations; the M5 Ultra has 4× more memory and draws 82% less power under load. Compare model compatibility, like-for-like tokens/sec and ownership cost below.

Dual RTX 5090 workstation (2x 32 GB) · Mac Studio M5 Ultra 256 GB

Verdict
  • throughputThe Dual RTX 5090 workstation (2x 32 GB) delivers approximately 57% higher aggregate decode throughput across 7 like-for-like model configurations (same quantization, runtime, context and inference settings).
  • capacityBoth machines run all 8 compared models in their best practical configurations. Compare their quantizations and evidence labels before treating any throughput difference as a hardware result.
  • costAt 8 hours a day, the Dual RTX 5090 workstation (2x 32 GB) has approximately 12% lower total monthly cost ($80 against $90), for a C corporation in United States (federal).
  • powerThe Dual RTX 5090 workstation (2x 32 GB) draws 1180 W under load against 215 W for the Mac Studio M5 Ultra 256 GB — a 5.5× difference that shows up in the electricity line and in how loud and hot the room gets.
  • evidenceAll 7 like-for-like throughput rows are estimates. Treat the ranking as indicative and the margins as uncertain.
  • recommendationThe Dual RTX 5090 workstation (2x 32 GB) is both faster and cheaper to own here, so it is the straightforward choice unless you need something specific to the Mac Studio M5 Ultra 256 GB.
Hardware
Dual RTX 5090 workstation (2x 32 GB)Mac Studio M5 Ultra 256 GB
Purchase price estimate$5,854 *$8,647 *
Memory64 GB256 GB
Theoretical bandwidth (aggregate on multi-GPU)3,584 GB/s1,200 GB/s
Usable bandwidth indicatorinterconnect-adjusted per workload1,052 GB/s (88%)
Nominal power1400 W300 W
Measured load power1180 W215 W
Idle power110 W13 W
Architecture2× GPUunified memory
Like-for-like performance7 comparable models
Ternary Bonsai 2 27BNemotron 3.5 Lightning 30B-A3BQwen3-Coder 30B-A3BQwen3 30B-A3BGemma 4 26B-A4Bgpt-oss-20bLlama 3.3 70BDeepSeek-R1-Distill 32B
ModelDual RTX 5090 workstation (2x 32 GB)Mac Studio M5 Ultra 256 GBWinner
Ternary Bonsai 2 27B
27.36B
no matched configuration
no matched configuration
Not comparable
Nemotron 3.5 Lightning 30B-A3B
31.578B (3B active)
~249–359 t/s
Q8_0
~180–258 t/s
Q8_0
Dual RTX 5090 workstation (2x 32 GB) +39%
estimated · estimator_v1 · 8K context
Qwen3-Coder 30B-A3B
30.532B (3.3B active)
~236–339 t/s
Q8_0
~167–240 t/s
Q8_0
Dual RTX 5090 workstation (2x 32 GB) +42%
estimated · estimator_v1 · 8K context
Qwen3 30B-A3B
30.532B (3.3B active)
~236–339 t/s
Q8_0
~167–240 t/s
Q8_0
Dual RTX 5090 workstation (2x 32 GB) +42%
estimated · estimator_v1 · 8K context
Gemma 4 26B-A4B
25.806B (3.8B active)
~139–200 t/s
BF16
~86.3–124 t/s
BF16
Dual RTX 5090 workstation (2x 32 GB) +61%
estimated · estimator_v1 · 8K context
gpt-oss-20b
20.915B (3.6B active)
~145–208 t/s
BF16
~90.6–130 t/s
BF16
Dual RTX 5090 workstation (2x 32 GB) +60%
estimated · estimator_v1 · 8K context
Llama 3.3 70B
70.554B
~30.8–44.3 t/s
Q5_K_M
~16.8–24.2 t/s
Q5_K_M
Dual RTX 5090 workstation (2x 32 GB) +83%
estimated · estimator_v1 · 8K context
DeepSeek-R1-Distill 32B
32.764B
~43.3–62.3 t/s
Q8_0
~24–34.6 t/s
Q8_0
Dual RTX 5090 workstation (2x 32 GB) +80%
estimated · estimator_v1 · 8K context
Hardware deltas are calculated only when quantization, runtime, exact context, execution/decode mode and relevant inference settings match. Measured and estimated results are never compared to each other. A difference under 10% is reported as a tie.
Best practical configuration
ModelDual RTX 5090 workstation (2x 32 GB)Mac Studio M5 Ultra 256 GBPractical result
Ternary Bonsai 2 27B
PTQ1_0 · estimated
PTQ1_0 · estimated
Both fit; configurations may differ
Nemotron 3.5 Lightning 30B-A3B~249–359 t/s
Q8_0 · estimated
~180–258 t/s
Q8_0 · estimated
Both fit; configurations may differ
Qwen3-Coder 30B-A3B~236–339 t/s
Q8_0 · estimated
~167–240 t/s
Q8_0 · estimated
Both fit; configurations may differ
Qwen3 30B-A3B~236–339 t/s
Q8_0 · estimated
~167–240 t/s
Q8_0 · estimated
Both fit; configurations may differ
Gemma 4 26B-A4B~216–311 t/s
Q8_0 · estimated
~149–214 t/s
Q8_0 · estimated
Both fit; configurations may differ
gpt-oss-20b~324–466 t/s
MXFP4 · estimated
~261–376 t/s
MXFP4 · estimated
Both fit; configurations may differ
Llama 3.3 70B~35.8–51.5 t/s
Q4_K_M · estimated
~11.3–16.3 t/s
Q8_0 · estimated
Both fit; configurations may differ
DeepSeek-R1-Distill 32B~43.3–62.3 t/s
Q8_0 · estimated
~24–34.6 t/s
Q8_0 · estimated
Both fit; configurations may differ
This is a deployment comparison: each machine uses its own recommended configuration. Different quantizations can change speed and quality, so these figures are not raw hardware benchmarks.
Economics — United States (federal), C corporation, 8h/day
Dual RTX 5090 workstation (2x 32 GB)Mac Studio M5 Ultra 256 GB
Purchase price estimate$5,854$8,647
Reclaimable VAT$0$0
Net cash at purchase$5,854$8,647
Monthly depreciation$82.94$100.88
Monthly electricity$7.98 (57 kWh)$1.32 (9 kWh)
Tax programme§179 100.0%§179 100.0%
Tax-programme benefit$1,229$1,816
Monthly economic cost$79.98$90.39
Total over 5 years$4,799$5,424
Purchase prices are localized inputs in USD: tax/VAT excluded for Dual RTX 5090 workstation (2x 32 GB) and tax/VAT excluded for Mac Studio M5 Ultra 256 GB. Net cash and monthly economic cost then apply the selected United States (federal) / C corporation tax treatment. Tax assumptions last verified 2026-09-07. Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.