Mac Studio M5 Max 64 GB
Apple's 2026 desktop workstation. The M5 generation adds Neural Accelerators to every GPU core, which lifts prefill throughput dramatically over M4 while decode stays governed by memory bandwidth. Its combination of very large unified memory and low idle power makes it the default choice for running large MoE models at home.
Applesystem15 models fit0 measured
Memory
64 GB
Official spec LPDDR5X unified
Memory bandwidth
614 GB/s
Official spec ~88% achieved in practice
Specifications
CPU
Apple M5 Max, 18-core Official spec
GPU
40-core GPU with Neural Accelerators Official spec
Memory
64 GB LPDDR5X unified Official spec
Memory bandwidth
614 GB/s Official spec source
Achieved bandwidth ⓘ
~88% (538 GB/s) Assumption
FP16 compute ⓘ
~120 TFLOPS Estimated
Architecture
Unified memory (CPU and GPU share one pool) Official spec
Nominal power
175 W Official spec
Measured load power
118 W Measured
Idle power
8 W Measured
Released
22 Sept 2026 Official spec
40-core GPU with 1 TB SSD, as listed on Apple's Dutch store. Price is ESTIMATED from the US-to-EUR 1.20x convention; Apple does not publish a per-option price list we can cite. Achieved-bandwidth coefficient INHERITED from mac-studio-m5-max-128gb, which shares the same silicon and software stack.
Standout models on this machine
Fastest
gpt-oss-20b — ~158–228 t/s
MXFP4 · estimated confidence
Largest that fits
Compare against
Related hardware
Model performance on Mac Studio M5 Max 64 GB0 measured, 15 estimated
15 of 15 rows
| Model↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|
| gpt-oss-20b OpenAI · 20.915B (3.6B active) | MXFP4 | 13.0 GB | ~158–228 t/s | ~1670–3470 t/s | 128K | Comfortable | Estimated |
| Qwen3-Coder 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~94.6–136 t/s | ~1450–3000 t/s | 128K | Comfortable | Estimated |
| Qwen3 30B-A3B Alibaba Qwen · 30.532B (3.3B active) | Q8_0 | 33.5 GB | ~94.6–136 t/s | ~1450–3000 t/s | 40K | Comfortable | Estimated |
| Gemma 4 26B-A4B Google DeepMind · 25.806B (3.8B active) | Q8_0 | 29.7 GB | ~83.4–120 t/s | ~1470–3040 t/s | 64K | Comfortable | Estimated |
| Qwen3 8B Alibaba Qwen · 8.191B | Q8_0 | 10.6 GB | ~47.7–68.7 t/s | ~1770–3680 t/s | 40K | Comfortable | Estimated |
| Phi-4 14B Microsoft · 14.66B | Q8_0 | 17.8 GB | ~27.4–39.4 t/s | ~990–2060 t/s | 16K | Comfortable | Estimated |
| Devstral Small 24B Mistral AI · 23.572B | Q8_0 | 26.8 GB | ~17.2–24.8 t/s | ~616–1280 t/s | 128K | Comfortable | Estimated |
| Mistral Small 3.2 24B Mistral AI · 24.011B | Q8_0 | 27.2 GB | ~16.9–24.4 t/s | ~604–1260 t/s | 128K | Comfortable | Estimated |
| Gemma 3 27B Google DeepMind · 27.432B | Q8_0 | 33.4 GB | ~14.9–21.4 t/s | ~529–1100 t/s | 32K | Comfortable | Estimated |
| Qwen3.8 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~14.7–21.1 t/s | ~522–1080 t/s | 64K | Comfortable | Estimated |
| Qwen3.6 27B Alibaba Qwen · 27.781B | Q8_0 | 31.9 GB | ~14.7–21.1 t/s | ~522–1080 t/s | 64K | Comfortable | Estimated |
| Gemma 4 31B Google DeepMind · 31.273B | Q8_0 | 41.0 GB | ~13.1–18.8 t/s | ~464–964 t/s | 16K | Comfortable | Estimated |
| Qwen3 32B Alibaba Qwen · 32.762B | Q8_0 | 37.1 GB | ~12.5–18 t/s | ~443–920 t/s | 32K | Comfortable | Estimated |
| DeepSeek-R1-Distill 32B DeepSeek · 32.764B | Q8_0 | 37.1 GB | ~12.5–18 t/s | ~443–920 t/s | 32K | Comfortable | Estimated |
| Llama 3.3 70B Meta · 70.554B | MLX 4-bit | 43.1 GB | ~11–15.8 t/s | ~206–427 t/s | 16K | Comfortable | Estimated |
Rows in grey are estimates from our bandwidth model, not measurements — they are shown as a range and never as a precise figure. Use the “measured only” filter to see just the 0 pairings on this machine that a real benchmark backs.
Ownership economicsUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$37
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $63 less
Net cash at purchase
$3,512
Calculated VAT not reclaimable
Total over 5 years
$2,213
Calculated after tax, after resale
Cost per USD/1M tokens
$1.51
Calculated 24.5M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Monthly breakdown
| Depreciation 2,459 over 5 years, straight-line to a 1,054 residual | $40.98 |
| Electricity 5.3 kWh/month at 0.140/kWh | $0.74 |
| Cost of capital 4.0%/yr on 2,283 average capital employed | $7.61 |
| Monthly cost before tax | $49.33 |
| Electricity tax shield Running costs are deductible business expenses | −$0.16 |
| First-year expensing §179 (100.0%) | −$12.29 |
| Monthly economic cost after tax | $36.88 |
Three different numbers, deliberately
Cash cost
$3,512
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$40.98/month
$2,459 written down over 5 years to a $1,054 residual.
After-tax economic cost
$36.88/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$1,054 (30%) Assumption
Extrapolated from the three-year curve.
Electricity
$0.140/kWh Assumption
Average power draw
30 W Calculated
Load 118 W for 20% of powered hours, idle 8 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $738 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$7.61/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.