Granite 4.2 30B

IBM's flagship Granite 4.2 reasoning model is a dense Apache-2.0 30B with switchable thinking modes, tool calling, multilingual support, and a native 128K context. Its 8 KV heads make it a straightforward fit-calculator reference for modern GQA.

IBMDensereasoningreasoning26 machines can run it
Run locally →
Total parameters
29.277B
Official spec sets memory need
Active parameters
29.277B
Official spec sets decode speed
Context
128K
Official spec tokens
4-bit weights
16 GB
Calculated before KV cache
Model specification
Publisher
IBM Official spec
Architecture
Dense transformer Official spec
Total parameters
29.277B Official spec
Context length
131,072 tokens Official spec
Attention
GQA · 8 KV heads × 128 dim × 64 layers Official spec
Licence
Apache 2.0 Official spec source
Specification confidence
High · verified 2026-09-21 Official spec source
Released
25 Aug 2026 Official spec
Official source
Quantizations and memory
QuantizationFormatBits/weightResidentDownloadQuality kept
MLX 4-bitmlx4.516.0 GB16.0 GB98.0%
Q4_K_Mdefaultgguf4.8517.0 GB17.0 GB98.5%
Q5_K_Mgguf5.6920.0 GB20.0 GB99.3%
Q8_0gguf8.529.5 GB29.5 GB99.9%
BF16safetensors1654.5 GB54.5 GB100.0%
Generic weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead. Runtime-specific artifacts use their published resident and download footprints; streamed models can therefore require much more disk than memory. Quality retention is an assumption, not a measured evaluation.
Coding & quality benchmarksCompare coding results →

No published benchmark results have been imported for this model yet. This is missing evidence, not a score of zero.

We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
~1.2–1.7 t/s · $1,531
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~31.4–45.2 t/s · $1,981
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~35.9–51.7 t/s · $2,071
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per purchase-price unit
~61.6–88.7 t/s · $3,332
Highest decode tokens/second per 1,000 units of the displayed purchase currency. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Granite 4.2 30B0 measured, 26 estimated
26 of 26 rows
HardwareQuantMemoryDecodePrefillContextPriceFitConfidence
NVIDIA DGX H200 (8x H200, 1,128 GB)
NVIDIA · 1,128 GB · 38,400 GB/s
Q8_0
recommended
36.2 GB~323–465 t/s~91390–189810 t/s128K$359,430ComfortableEstimated
Quad RTX 5090 workstation (4x 32 GB)
NVIDIA · 128 GB · 7,168 GB/s
Q8_0
recommended
34.6 GB~72.1–104 t/s~3420–7110 t/s128K$12,611ComfortableEstimated
Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB)
Lenovo · 384 GB · 7,168 GB/s
Q8_0
recommended
34.6 GB~63.1–90.7 t/s~4390–9110 t/s128K$58,553ComfortableEstimated
RTX 5090 workstation (1x 32 GB)
NVIDIA · 32 GB · 1,792 GB/s
Q4_K_M
recommended
20.5 GB~61.6–88.7 t/s~1050–2190 t/s32K$3,332ComfortableEstimated
Dual RTX 5090 workstation (2x 32 GB)
NVIDIA · 64 GB · 3,584 GB/s
Q8_0
recommended
33.8 GB~48.1–69.2 t/s~1710–3550 t/s64K$5,854ComfortableEstimated
Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB)
Dell · 192 GB · 3,584 GB/s
Q8_0
recommended
33.8 GB~44–63.4 t/s~2190–4560 t/s128K$36,663ComfortableEstimated
Dual used RTX 4090 workstation (48 GB)
NVIDIA · 48 GB · 2,016 GB/s
Q5_K_M
recommended
24.0 GB~40.4–58.1 t/s~1900–3960 t/s64K$3,692ComfortableEstimated
Used RTX 4090 workstation (24 GB)
NVIDIA · 24 GB · 1,008 GB/s
Q4_K_M
recommended
20.5 GB~35.9–51.7 t/s~1170–2430 t/s8K$2,071BorderlineEstimated
RTX PRO 6000 Blackwell workstation (96 GB)
NVIDIA · 96 GB · 1,792 GB/s
Q8_0
recommended
33.4 GB~33.9–48.8 t/s~1110–2310 t/s128K$10,359ComfortableEstimated
RTX PRO 6000 Max-Q workstation (96 GB, 300 W)
NVIDIA · 96 GB · 1,792 GB/s
Q8_0
recommended
33.4 GB~33.9–48.8 t/s~910–1890 t/s128K$10,809ComfortableEstimated
RTX 3090 workstation (Ryzen 9 7950X, 64 GB)
NVIDIA · 24 GB · 936 GB/s
Q4_K_M
recommended
20.5 GB~31.4–45.2 t/s~504–1050 t/s8K$1,981BorderlineEstimated
Radeon AI PRO R9700 workstation (32 GB)
AMD · 32 GB · 644 GB/s
Q4_K_M
recommended
20.5 GB~27.5–39.6 t/s~839–1740 t/s32K$2,521ComfortableEstimated
NVIDIA DGX Station GB300 (748 GB)
NVIDIA · 748 GB · 7,100 GB/s
Q8_0
recommended
33.4 GB~27.1–39.1 t/s~12210–25360 t/s128K$90,082ComfortableEstimated
Mac Studio M5 Ultra 512 GB
Apple · 512 GB · 1,200 GB/s
Q8_0
recommended
33.4 GB~26.8–38.6 t/s~991–2060 t/s128K$11,350ComfortableEstimated
Mac Studio M5 Ultra 256 GB
Apple · 256 GB · 1,200 GB/s
Q8_0
recommended
33.4 GB~26.8–38.6 t/s~991–2060 t/s128K$8,647ComfortableEstimated
Mac Studio M5 Ultra 96 GB
Apple · 96 GB · 1,200 GB/s
Q8_0
recommended
33.4 GB~26.8–38.6 t/s~793–1650 t/s128K$5,945ComfortableEstimated
Mac Studio M5 Max 36 GB
Apple · 36 GB · 614 GB/s
Q5_K_M
recommended
23.6 GB~20.7–29.7 t/s~397–824 t/s32K$2,702ComfortableEstimated
Mac Studio M3 Ultra 512 GB
Apple · 512 GB · 819 GB/s
Q8_0
recommended
33.4 GB~15–21.6 t/s~239–497 t/s128K$8,287ComfortableEstimated
Mac Studio M3 Ultra 256 GB
Apple · 256 GB · 819 GB/s
Q8_0
recommended
33.4 GB~15–21.6 t/s~239–497 t/s128K$5,674ComfortableEstimated
Mac Studio M5 Max 128 GB
Apple · 128 GB · 614 GB/s
Q8_0
recommended
33.4 GB~13.9–20.1 t/s~496–1030 t/s128K$4,323ComfortableEstimated
Mac Studio M5 Max 64 GB
Apple · 64 GB · 614 GB/s
Q8_0
recommended
33.4 GB~13.9–20.1 t/s~496–1030 t/s64K$3,512ComfortableEstimated
Mac mini M4 Pro 64 GB
Apple · 64 GB · 273 GB/s
Q8_0
recommended
33.4 GB~5–7.2 t/s~41.5–86.2 t/s64K$2,251ComfortableEstimated
NVIDIA DGX Spark 128 GB
NVIDIA · 128 GB · 273 GB/s
Q8_0
recommended
33.4 GB~4.8–6.9 t/s~427–886 t/s128K$3,873ComfortableEstimated
Framework Desktop (Ryzen AI Max+ 395, 128 GB)
AMD · 128 GB · 256 GB/s
Q8_0
recommended
33.4 GB~3.2–4.7 t/s~118–245 t/s128K$2,161ComfortableEstimated
GMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB)
AMD · 128 GB · 256 GB/s
Q8_0
recommended
33.4 GB~2.8–4 t/s~118–245 t/s128K$1,711ComfortableEstimated
CPU-only workstation (Ryzen 9950X, 192 GB DDR5)
Generic · 192 GB · 90 GB/s
Q8_0
recommended
33.4 GB~1.2–1.7 t/s~14.2–29.5 t/s128K$1,531ComfortableEstimated
Submit a benchmarkContributions are reviewed before publication.
Similar models