Open-weight models
27 models with 96 quantized variants. Two numbers decide whether a machine can run a model well: the total parameter count sets how much memory you need, and the active parameter count sets how fast it generates. A sparse mixture of experts can be enormous and still fast.
All models (27)
| Model | Publisher | Type | Total | Active | Context | 4-bit size | Licence | Index |
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4 Pro reasoning · reasoning | DeepSeek | MoE | 1.6T | 49B | 1M | 871 GB | MIT | — |
| Kimi K2.6 agentic · reasoning | Moonshot AI | MoE | 1.0T | 32B | 256K | 559 GB | Modified MIT | — |
| GLM-5.3 coding · reasoning | Z.ai | MoE | 753.33B | 40B | 1M | 410 GB | MIT | — |
| DeepSeek-V3.2 general · reasoning | DeepSeek | MoE | 685.397B | 37B | 160K | 373 GB | MIT | — |
| MiniMax M3 agentic · reasoning | MiniMax | MoE | 427.04B | 23B | 1M | 233 GB | Apache 2.0 | — |
| GLM-5.3 Flash coding · reasoning | Z.ai | MoE | 321.323B | 18B | 1M | 175 GB | MIT | — |
| DeepSeek-V4 Flash coding · reasoning | DeepSeek | MoE | 304.18B | 13B | 1M | 166 GB | MIT | — |
| Hunyuan Hy3 general · reasoning | Tencent | MoE | 298.786B | 21B | 256K | 163 GB | Tencent Hunyuan Community | — |
| Qwen3 235B-A22B general · reasoning | Alibaba Qwen | MoE | 235.094B | 22B | 256K | 128 GB | Apache 2.0 | — |
| Qwen3.8 Flash Next agentic · reasoning | Alibaba Qwen | MoE | 180B | 6B | 256K | 98 GB | Qwen Community 1.0 | — |
| Qwen3.5 122B-A10B general · reasoning | Alibaba Qwen | MoE | 125.086B | 10B | 256K | 68 GB | Apache 2.0 | — |
| gpt-oss-120b general · reasoning | OpenAI | MoE | 116.829B | 5.1B | 128K | 63 GB | Apache 2.0 | — |
| Llama 3.3 70B general | Meta | Dense | 70.554B | — | 128K | 38 GB | Llama 3.3 Community License | — |
| DeepSeek-R1-Distill 32B reasoning · reasoning | DeepSeek | Dense | 32.764B | — | 128K | 18 GB | MIT | — |
| Qwen3 32B general · reasoning | Alibaba Qwen | Dense | 32.762B | — | 40K | 18 GB | Apache 2.0 | — |
| Gemma 4 31B general | Google DeepMind | Dense | 31.273B | — | 256K | 17 GB | Gemma Terms of Use | — |
| Qwen3-Coder 30B-A3B coding | Alibaba Qwen | MoE | 30.532B | 3.3B | 256K | 17 GB | Apache 2.0 | — |
| Qwen3 30B-A3B general · reasoning | Alibaba Qwen | MoE | 30.532B | 3.3B | 40K | 17 GB | Apache 2.0 | — |
| Qwen3.8 27B general · reasoning | Alibaba Qwen | Dense | 27.781B | — | 256K | 15 GB | Apache 2.0 | — |
| Qwen3.6 27B coding | Alibaba Qwen | Dense | 27.781B | — | 256K | 15 GB | Apache 2.0 | — |
| Gemma 3 27B general | Google DeepMind | Dense | 27.432B | — | 128K | 15 GB | Gemma Terms of Use | — |
| Gemma 4 26B-A4B general | Google DeepMind | MoE | 25.806B | 3.8B | 256K | 14 GB | Gemma Terms of Use | — |
| Mistral Small 3.2 24B general | Mistral AI | Dense | 24.011B | — | 128K | 13 GB | Apache 2.0 | — |
| Devstral Small 24B coding | Mistral AI | Dense | 23.572B | — | 128K | 13 GB | Apache 2.0 | — |
| gpt-oss-20b general · reasoning | OpenAI | MoE | 20.915B | 3.6B | 128K | 11 GB | Apache 2.0 | — |
| Phi-4 14B reasoning | Microsoft | Dense | 14.66B | — | 16K | 8 GB | MIT | — |
| Qwen3 8B general · reasoning | Alibaba Qwen | Dense | 8.191B | — | 40K | 4 GB | Apache 2.0 | — |
“4-bit size” is the computed weight footprint of the default quantization, before the KV cache and runtime overhead — the model page for each shows the full memory requirement on real hardware. “Index” is the BenchLM aggregate open-weight score where published; it is one aggregator’s view, not a universal quality number, and we deliberately do not compute one of our own.