MiMo-V2.6 Flash 309B-A15B

XiaomiMiMo's efficiency-focused MIT-licensed omnimodal MoE has 309B total and 15B active parameters, plus a one-million-token context. It handles text, images, video, and audio, with 256 routed experts and 8 active per token.

XiaomiMiMoMixture of expertsagenticreasoning7 machines can run it
Total parameters
309B
Official spec sets memory need
Active parameters
15B
Official spec sets decode speed
Context
1M
Official spec tokens
4-bit weights
168 GB
Calculated before KV cache
This is a sparse mixture of experts. All 309B of weights must be available to the runtime, but only 15B are read per generated token. Ordinary runtimes keep the full quantized model in memory; a runtime-specific SSD-streaming artifact can retain a smaller working set and fetch expert data on demand, trading speed for capacity.
Model specification
Publisher
XiaomiMiMo Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
309B Official spec
Active parameters
15B per token Official spec
Context length
1,048,576 tokens Official spec
Attention
GQA · 4 KV heads × 192 dim × 48 layers Official spec
Licence
MIT Official spec source
Specification confidence
High · verified 2026-09-22 Official spec source
Released
22 Sept 2026 Official spec
Official source
Quantizations and memory
QuantizationFormatBits/weightResidentDownloadQuality kept
MLX 4-bitmlx4.5168.4 GB168.4 GB98.0%
Q4_K_Mdefaultgguf4.85179.7 GB179.7 GB98.5%
Q5_K_Mgguf5.69210.8 GB210.8 GB99.3%
Q8_0gguf8.5311.9 GB311.9 GB99.9%
Generic weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead. Runtime-specific artifacts use their published resident and download footprints; streamed models can therefore require much more disk than memory. Quality retention is an assumption, not a measured evaluation.
Coding & quality benchmarksCompare coding results →

No published benchmark results have been imported for this model yet. This is missing evidence, not a score of zero.

We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
~42.2–60.7 t/s · $5,674
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~42.2–60.7 t/s · $5,674
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~42.2–60.7 t/s · $5,674
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per purchase-price unit
~73.4–106 t/s · $8,647
Highest decode tokens/second per 1,000 units of the displayed purchase currency. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs MiMo-V2.6 Flash 309B-A15B0 measured, 7 estimated
7 of 7 rows
HardwareQuantMemoryDecodePrefillContextPriceFitConfidence
NVIDIA DGX H200 (8x H200, 1,128 GB)
NVIDIA · 1,128 GB · 38,400 GB/s
Q8_0
recommended
326.2 GB~393–566 t/s~39300–81620 t/s1024K$359,430ComfortableEstimated
Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB)
Lenovo · 384 GB · 7,168 GB/s
Q5_K_M
recommended
220.5 GB~135–194 t/s~1890–3920 t/s512K$58,553ComfortableEstimated
Mac Studio M5 Ultra 256 GB
Apple · 256 GB · 1,200 GB/s
Q4_K_M
recommended
187.2 GB~73.4–106 t/s~426–885 t/s128K$8,647FitsEstimated
NVIDIA DGX Station GB300 (748 GB)
NVIDIA · 748 GB · 7,100 GB/s
Q8_0
recommended
323.4 GB~44.1–63.4 t/s~5250–10910 t/s1024K$90,082ComfortableEstimated
Mac Studio M5 Ultra 512 GB
Apple · 512 GB · 1,200 GB/s
Q8_0
recommended
323.4 GB~43.5–62.7 t/s~426–885 t/s512K$11,350ComfortableEstimated
Mac Studio M3 Ultra 256 GB
Apple · 256 GB · 819 GB/s
Q4_K_M
recommended
187.2 GB~42.2–60.7 t/s~103–214 t/s128K$5,674FitsEstimated
Mac Studio M3 Ultra 512 GB
Apple · 512 GB · 819 GB/s
Q8_0
recommended
323.4 GB~24.6–35.4 t/s~103–214 t/s512K$8,287ComfortableEstimated
Submit a benchmarkContributions are reviewed before publication.
Similar models