Kimi K2.6
A trillion-parameter agentic MoE with 32B active. Needs 512 GB or more even at 4-bit. Included because it is one of the strongest open-weight models for tool use, and because the fit calculator's answer for most hardware here is a clear no.
Moonshot AIMixture of expertsagenticreasoning0 machines can run it
Total parameters
1.0T
Official spec sets memory need
Active parameters
32B
Official spec sets decode speed
Context
256K
Official spec tokens
4-bit weights
559 GB
Calculated before KV cache
This is a sparse mixture of experts. All 1.0T of weights must sit in memory, but only 32B are read per generated token — so it needs the memory of a 1.0T model and generates at roughly the speed of a 32B one. That is why large unified-memory machines suit it and fast 32 GB GPUs do not.
Model specification
Publisher
Moonshot AI Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
1.0T Official spec
Active parameters
32B per token Official spec
Context length
262,144 tokens Official spec
Attention ⓘ
MLA · 64 KV heads × 192 dim × 61 layers Official spec
Licence
Modified MIT Official spec source
Specification confidence ⓘ
High · verified 2026-09-07 Official spec source
Released
1 Jun 2026 Official spec
Official source
Quantizations and memory
| Quantization | Format | Bits/weight | Weights | Quality kept |
|---|---|---|---|---|
MLX 4-bit | mlx | 4.5 | 559.5 GB | 98.0% |
Q4_K_Mdefault | gguf | 4.85 | 597.2 GB | 98.5% |
Q8_0 | gguf | 8.5 | 1036.5 GB | 99.9% |
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Quality benchmarks
| Benchmark | Category | Score | Reported by | Date | Source |
|---|---|---|---|---|---|
| SWE-bench Pro | coding | 58.6 | canonical model_repository | 20 Apr 2026 | link |
| SWE-bench Verified | coding | 80.2 | canonical model_repository | 20 Apr 2026 | link |
| GPQA | reasoning | 90.5 | canonical model_repository | 20 Apr 2026 | link |
We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
Nothing in the database qualifies.
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
Nothing in the database qualifies.
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
Nothing in the database qualifies.
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per euro
Nothing in the database qualifies.
Highest decode tokens/second per EUR 1,000 of purchase price. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Kimi K2.60 measured, 0 estimated
0 of 0 rows
| Hardware↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Price↕ | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|---|
| Nothing matches these filters. | ||||||||
Similar models
MiniMax M3
MiniMax · 427.04B (23B active) · Apache 2.0
Qwen3.8 Flash Next
Alibaba Qwen · 180B (6B active) · Qwen Community 1.0
GLM-5.3
Z.ai · 753.33B (40B active) · MIT
DeepSeek-V3.2
DeepSeek · 685.397B (37B active) · MIT
DeepSeek-V4 Pro
DeepSeek · 1.6T (49B active) · MIT
GLM-5.3 Flash
Z.ai · 321.323B (18B active) · MIT