Open-weight models

27 models with 96 quantized variants. Two numbers decide whether a machine can run a model well: the total parameter count sets how much memory you need, and the active parameter count sets how fast it generates. A sparse mixture of experts can be enormous and still fast.

All models (27)
ModelPublisherTypeTotalActiveContext4-bit sizeLicenceIndex
DeepSeek-V4 Pro
reasoning · reasoning
DeepSeekMoE1.6T49B1M871 GBMIT
Kimi K2.6
agentic · reasoning
Moonshot AIMoE1.0T32B256K559 GBModified MIT
GLM-5.3
coding · reasoning
Z.aiMoE753.33B40B1M410 GBMIT
DeepSeek-V3.2
general · reasoning
DeepSeekMoE685.397B37B160K373 GBMIT
MiniMax M3
agentic · reasoning
MiniMaxMoE427.04B23B1M233 GBApache 2.0
GLM-5.3 Flash
coding · reasoning
Z.aiMoE321.323B18B1M175 GBMIT
DeepSeek-V4 Flash
coding · reasoning
DeepSeekMoE304.18B13B1M166 GBMIT
Hunyuan Hy3
general · reasoning
TencentMoE298.786B21B256K163 GBTencent Hunyuan Community
Qwen3 235B-A22B
general · reasoning
Alibaba QwenMoE235.094B22B256K128 GBApache 2.0
Qwen3.8 Flash Next
agentic · reasoning
Alibaba QwenMoE180B6B256K98 GBQwen Community 1.0
Qwen3.5 122B-A10B
general · reasoning
Alibaba QwenMoE125.086B10B256K68 GBApache 2.0
gpt-oss-120b
general · reasoning
OpenAIMoE116.829B5.1B128K63 GBApache 2.0
Llama 3.3 70B
general
MetaDense70.554B128K38 GBLlama 3.3 Community License
DeepSeek-R1-Distill 32B
reasoning · reasoning
DeepSeekDense32.764B128K18 GBMIT
Qwen3 32B
general · reasoning
Alibaba QwenDense32.762B40K18 GBApache 2.0
Gemma 4 31B
general
Google DeepMindDense31.273B256K17 GBGemma Terms of Use
Qwen3-Coder 30B-A3B
coding
Alibaba QwenMoE30.532B3.3B256K17 GBApache 2.0
Qwen3 30B-A3B
general · reasoning
Alibaba QwenMoE30.532B3.3B40K17 GBApache 2.0
Qwen3.8 27B
general · reasoning
Alibaba QwenDense27.781B256K15 GBApache 2.0
Qwen3.6 27B
coding
Alibaba QwenDense27.781B256K15 GBApache 2.0
Gemma 3 27B
general
Google DeepMindDense27.432B128K15 GBGemma Terms of Use
Gemma 4 26B-A4B
general
Google DeepMindMoE25.806B3.8B256K14 GBGemma Terms of Use
Mistral Small 3.2 24B
general
Mistral AIDense24.011B128K13 GBApache 2.0
Devstral Small 24B
coding
Mistral AIDense23.572B128K13 GBApache 2.0
gpt-oss-20b
general · reasoning
OpenAIMoE20.915B3.6B128K11 GBApache 2.0
Phi-4 14B
reasoning
MicrosoftDense14.66B16K8 GBMIT
Qwen3 8B
general · reasoning
Alibaba QwenDense8.191B40K4 GBApache 2.0

“4-bit size” is the computed weight footprint of the default quantization, before the KV cache and runtime overhead — the model page for each shows the full memory requirement on real hardware. “Index” is the BenchLM aggregate open-weight score where published; it is one aggregator’s view, not a universal quality number, and we deliberately do not compute one of our own.