gpt-oss-20b
The small MXFP4 sibling: about 12 GB of weights and 3.6B active parameters. Fast on essentially every machine in this database, including CPU-only builds.
OpenAIMixture of expertsgeneralreasoning21 machines can run it
Total parameters
20.915B
Official spec sets memory need
Active parameters
3.6B
Official spec sets decode speed
Context
128K
Official spec tokens
4-bit weights
11 GB
Calculated before KV cache
This is a sparse mixture of experts. All 20.915B of weights must sit in memory, but only 3.6B are read per generated token — so it needs the memory of a 20.915B model and generates at roughly the speed of a 3.6B one. That is why large unified-memory machines suit it and fast 32 GB GPUs do not.
Model specification
Publisher
OpenAI Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
20.915B Official spec
Active parameters
3.6B per token Official spec
Context length
131,072 tokens Official spec
Attention ⓘ
GQA · 8 KV heads × 64 dim × 24 layers Official spec
Licence
Apache 2.0 Official spec source
Specification confidence ⓘ
High · verified 2026-09-07 Official spec source
Released
5 Aug 2025 Official spec
Official source
Quantizations and memory
| Quantization | Format | Bits/weight | Weights | Quality kept |
|---|---|---|---|---|
MXFP4default | gguf | 4.25 | 11.3 GB | 100.0% |
BF16 | safetensors | 16 | 39.0 GB | 100.0% |
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Recommendations
Cheapest that can run it
~16.1–23.1 t/s · €1.699
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~36.4–52.4 t/s · €1.899
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~36.4–52.4 t/s · €1.899
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
419 t/s · €3.699
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per euro
~228–329 t/s · €2.299
Highest decode tokens/second per EUR 1,000 of purchase price. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs gpt-oss-20b1 measured, 20 estimated
21 of 21 rows
| Hardware↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Price↕ | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|---|
| Quad RTX 5090 workstation (4x 32 GB) NVIDIA · 128 GB · 7,168 GB/s | MXFP4 | 14.2 GB | ~387–556 t/s | ~23090–47960 t/s | 128K | €13.999 | Comfortable | Estimated |
| RTX 5090 workstation (1x 32 GB) NVIDIA · 32 GB · 1,792 GB/s | MXFP4 | 13.0 GB | 419 t/s | — | 128K | €3.699 | Comfortable | Low |
| RTX PRO 6000 Blackwell workstation (96 GB) NVIDIA · 96 GB · 1,792 GB/s | MXFP4 | 13.0 GB | ~329–474 t/s | ~7500–15580 t/s | 128K | €11.499 | Comfortable | Estimated |
| RTX PRO 6000 Max-Q workstation (96 GB, 300 W) NVIDIA · 96 GB · 1,792 GB/s | MXFP4 | 13.0 GB | ~329–474 t/s | ~6140–12760 t/s | 128K | €11.999 | Comfortable | Estimated |
| Dual RTX 5090 workstation (2x 32 GB) NVIDIA · 64 GB · 3,584 GB/s | MXFP4 | 13.4 GB | ~324–466 t/s | ~11550–23980 t/s | 128K | €6.499 | Comfortable | Estimated |
| Mac Studio M5 Ultra 512 GB Apple · 512 GB · 1,200 GB/s | MXFP4 | 13.0 GB | ~261–376 t/s | ~3340–6950 t/s | 128K | €12.599 | Comfortable | Estimated |
| Mac Studio M5 Ultra 256 GB Apple · 256 GB · 1,200 GB/s | MXFP4 | 13.0 GB | ~261–376 t/s | ~3340–6950 t/s | 128K | €9.599 | Comfortable | Estimated |
| Mac Studio M5 Ultra 96 GB Apple · 96 GB · 1,200 GB/s | MXFP4 | 13.0 GB | ~261–376 t/s | ~2680–5560 t/s | 128K | €6.599 | Comfortable | Estimated |
| Dual used RTX 4090 workstation (48 GB) NVIDIA · 48 GB · 2,016 GB/s | MXFP4 | 13.4 GB | ~238–343 t/s | ~12850–26700 t/s | 128K | €4.099 | Comfortable | Estimated |
| Used RTX 4090 workstation (24 GB) NVIDIA · 24 GB · 1,008 GB/s | MXFP4 | 13.0 GB | ~228–329 t/s | ~7910–16430 t/s | 128K | €2.299 | Comfortable | Estimated |
| Radeon AI PRO R9700 workstation (32 GB) AMD · 32 GB · 644 GB/s | MXFP4 | 13.0 GB | ~183–264 t/s | ~2830–5880 t/s | 128K | €2.799 | Comfortable | Estimated |
| Mac Studio M3 Ultra 512 GB Apple · 512 GB · 819 GB/s | MXFP4 | 13.0 GB | ~168–242 t/s | ~807–1680 t/s | 128K | €9.199 | Comfortable | Estimated |
| Mac Studio M3 Ultra 256 GB Apple · 256 GB · 819 GB/s | MXFP4 | 13.0 GB | ~168–242 t/s | ~807–1680 t/s | 128K | €6.299 | Comfortable | Estimated |
| Mac Studio M5 Max 128 GB Apple · 128 GB · 614 GB/s | MXFP4 | 13.0 GB | ~158–228 t/s | ~1670–3470 t/s | 128K | €4.799 | Comfortable | Estimated |
| Mac Studio M5 Max 64 GB Apple · 64 GB · 614 GB/s | MXFP4 | 13.0 GB | ~158–228 t/s | ~1670–3470 t/s | 128K | €3.899 | Comfortable | Estimated |
| Mac Studio M5 Max 36 GB Apple · 36 GB · 614 GB/s | MXFP4 | 13.0 GB | ~158–228 t/s | ~1340–2780 t/s | 128K | €2.999 | Comfortable | Estimated |
| Mac mini M4 Pro 64 GB Apple · 64 GB · 273 GB/s | MXFP4 | 13.0 GB | ~64.2–92.4 t/s | ~140–291 t/s | 128K | €2.499 | Comfortable | Estimated |
| NVIDIA DGX Spark 128 GB NVIDIA · 128 GB · 273 GB/s | MXFP4 | 13.0 GB | ~59.4–85.5 t/s | ~2880–5980 t/s | 128K | €4.299 | Comfortable | Estimated |
| Framework Desktop (Ryzen AI Max+ 395, 128 GB) AMD · 128 GB · 256 GB/s | MXFP4 | 13.0 GB | ~41.6–59.8 t/s | ~398–826 t/s | 128K | €2.399 | Comfortable | Estimated |
| GMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB) AMD · 128 GB · 256 GB/s | MXFP4 | 13.0 GB | ~36.4–52.4 t/s | ~398–826 t/s | 128K | €1.899 | Comfortable | Estimated |
| CPU-only workstation (Ryzen 9950X, 192 GB DDR5) Generic · 192 GB · 90 GB/s | MXFP4 | 13.0 GB | ~16.1–23.1 t/s | ~47.9–99.6 t/s | 128K | €1.699 | Comfortable | Estimated |
Similar models
Gemma 4 26B-A4B
Google DeepMind · 25.806B (3.8B active) · Gemma Terms of Use
Qwen3 30B-A3B
Alibaba Qwen · 30.532B (3.3B active) · Apache 2.0
gpt-oss-120b
OpenAI · 116.829B (5.1B active) · Apache 2.0
Mistral Small 3.2 24B
Mistral AI · 24.011B · Apache 2.0
Gemma 3 27B
Google DeepMind · 27.432B · Gemma Terms of Use
Qwen3.8 27B
Alibaba Qwen · 27.781B · Apache 2.0