Best hardware for local LLMs by budget
New-system catalogue picks for Qwen3.8 27B at 8K context. US prices are planning estimates, not live offers.
Fastest Q4_K_M estimate in this band · best purchase-price efficiency too
$2,251 estimated US catalogue price · ~9.2–13.2 t/s decode at 8K context; compared with 2 other new configurations at the same model and quantization. Also 5.0 central-estimate t/s per $1,000 catalogue price. Estimated, not measured. Estimated
Memory-focused alternative
GMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB)
$1,711 estimated US catalogue price · 128 GB installed. Unified memory is shared with the system; not all of it is available to the model. ~5.2–7.4 t/s estimated decode at Q4_K_M. Estimated
Fastest Q4_K_M estimate in this band · best purchase-price efficiency too
RTX 5090 workstation (1x 32 GB)
$3,332 estimated US catalogue price · ~64.7–93.1 t/s decode at 8K context; compared with 5 other new configurations at the same model and quantization. Also 23.7 central-estimate t/s per $1,000 catalogue price. Estimated, not measured. Estimated
Memory-focused alternative
$4,323 estimated US catalogue price · 128 GB installed. Unified memory is shared with the system; not all of it is available to the model. ~25.4–36.5 t/s estimated decode at Q4_K_M. Estimated
AMD/ROCm alternative
Radeon AI PRO R9700 workstation (32 GB)
$2,521 estimated US catalogue price · 32 GB installed; ~28.9–41.7 t/s estimated decode at Q4_K_M. Check runtime support and model fit before choosing a GPU stack. Estimated
Fastest Q4_K_M estimate in this band · best purchase-price efficiency too
Dual RTX 5090 workstation (2x 32 GB)
$5,854 estimated US catalogue price · ~83–119 t/s decode at 8K context; compared with 2 other new configurations at the same model and quantization. Also 17.3 central-estimate t/s per $1,000 catalogue price. Estimated, not measured. Estimated
Memory-focused alternative
$8,647 estimated US catalogue price · 256 GB installed. Unified memory is shared with the system; not all of it is available to the model. ~48.2–69.3 t/s estimated decode at Q4_K_M. Estimated
A used RTX 4090 workstation can be worth comparing with these new systems. Used asking prices vary by region, condition and the rest of the build, so this guide does not rank them from a converted catalogue price. Inspect used-listing snapshots and their collection dates →
Start with the model you actually plan to run. These speed estimates use one quantization for a fair comparison; the best quantization for your machine may differ. Check model fit, runtime support and the price available where you live.