Qwen3.8 Flash Next
176B total but only 6B active — the most extreme sparsity ratio in this catalogue. That makes it unusually fast on low-bandwidth unified-memory machines like the DGX Spark and Strix Halo, where a dense model of comparable quality would be unusable.
Alibaba QwenMixture of expertsagenticreasoning7 machines can run it
Total parameters
180B
Official spec sets memory need
Active parameters
6B
Official spec sets decode speed
Context
256K
Official spec tokens
4-bit weights
98 GB
Calculated before KV cache
This is a sparse mixture of experts. All 180B of weights must sit in memory, but only 6B are read per generated token — so it needs the memory of a 180B model and generates at roughly the speed of a 6B one. That is why large unified-memory machines suit it and fast 32 GB GPUs do not.
Model specification
Publisher
Alibaba Qwen Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
180B Official spec
Active parameters
6B per token Official spec
Context length
262,144 tokens Official spec
Attention ⓘ
GQA · 2 KV heads × 256 dim × 48 layers Official spec
Licence
Qwen Community 1.0 Official spec source
Specification confidence ⓘ
High · verified 2026-09-07 Official spec source
Released
1 Aug 2026 Official spec
Official source
Quantizations and memory
| Quantization | Format | Bits/weight | Weights | Quality kept |
|---|---|---|---|---|
MLX 4-bit | mlx | 4.5 | 98.1 GB | 98.0% |
Q4_K_Mdefault | gguf | 4.85 | 104.7 GB | 98.5% |
Q8_0 | gguf | 8.5 | 181.7 GB | 99.9% |
BF16 | safetensors | 16 | 335.3 GB | 100.0% |
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Recommendations
Cheapest that can run it
~8.5–12.3 t/s · €1.699
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~97.9–141 t/s · €4.799
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~97.9–141 t/s · €4.799
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per euro
~307–442 t/s · €13.999
Highest decode tokens/second per EUR 1,000 of purchase price. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Qwen3.8 Flash Next0 measured, 7 estimated
7 of 7 rows
| Hardware↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Price↕ | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|---|
| Quad RTX 5090 workstation (4x 32 GB) NVIDIA · 128 GB · 7,168 GB/s | MLX 4-bit | 104.0 GB | ~307–442 t/s | ~6100–12660 t/s | — | €13.999 | Borderline | Estimated |
| Mac Studio M5 Ultra 256 GB Apple · 256 GB · 1,200 GB/s | Q4_K_M | 109.6 GB | ~162–233 t/s | ~883–1830 t/s | 256K | €9.599 | Comfortable | Estimated |
| Mac Studio M5 Ultra 512 GB Apple · 512 GB · 1,200 GB/s | Q8_0 | 188.9 GB | ~101–145 t/s | ~883–1830 t/s | 256K | €12.599 | Comfortable | Estimated |
| Mac Studio M3 Ultra 256 GB Apple · 256 GB · 819 GB/s | Q4_K_M | 109.6 GB | ~97.9–141 t/s | ~213–442 t/s | 256K | €6.299 | Comfortable | Estimated |
| Mac Studio M5 Max 128 GB Apple · 128 GB · 614 GB/s | MLX 4-bit | 102.8 GB | ~97.9–141 t/s | ~442–917 t/s | — | €4.799 | Borderline | Estimated |
| Mac Studio M3 Ultra 512 GB Apple · 512 GB · 819 GB/s | Q8_0 | 188.9 GB | ~58.9–84.7 t/s | ~213–442 t/s | 256K | €9.199 | Comfortable | Estimated |
| CPU-only workstation (Ryzen 9950X, 192 GB DDR5) Generic · 192 GB · 90 GB/s | Q4_K_M | 109.6 GB | ~8.5–12.3 t/s | ~12.7–26.3 t/s | 256K | €1.699 | Comfortable | Estimated |
Similar models
MiniMax M3
MiniMax · 427.04B (23B active) · Apache 2.0
Qwen3 235B-A22B
Alibaba Qwen · 235.094B (22B active) · Apache 2.0
Qwen3.5 122B-A10B
Alibaba Qwen · 125.086B (10B active) · Apache 2.0
Kimi K2.6
Moonshot AI · 1.0T (32B active) · Modified MIT
gpt-oss-120b
OpenAI · 116.829B (5.1B active) · Apache 2.0
Hunyuan Hy3
Tencent · 298.786B (21B active) · Tencent Hunyuan Community