Ternary Bonsai 2 27B
Prism ML's ternary derivative of Qwen3.8 27B. The text-only GGUF packs are 5.95 GB (PTQ1_0) and 7.21 GB (PQ2_0), with an optional 0.63 GB vision projector. A separate MLX pack includes the vision tower. Its Hadamard-rotated weights require Prism ML's custom runtimes; stock llama.cpp cannot run the ternary GGUF files. Publisher-reported quality and speed have not been independently verified here.
Prism MLDensereasoningreasoning2 machines can run it
Total parameters
27.36B
Official spec sets memory need
Active parameters
27.36B
Official spec sets decode speed
Context
256K
Official spec tokens
4-bit weights
6 GB
Calculated before KV cache
Model specification
Publisher
Prism ML Official spec
Architecture
Dense transformer Official spec
Total parameters
27.36B Official spec
Context length
262,144 tokens Official spec
Attention ⓘ
HYBRID_LINEAR · 4 KV heads × 256 dim × 64 layers Official spec
Licence
Apache 2.0 Official spec source
Specification confidence ⓘ
High · verified 2026-09-21 Official spec source
Released
17 Sept 2026 Official spec
Official source
Quantizations and memory
| Quantization | Format | Bits/weight | Resident | Download | Quality kept |
|---|---|---|---|---|---|
PTQ1_0default | gguf | 1.75 | 5.5 GB | 5.5 GB | — |
PQ2_0 | gguf | 2.13 | 6.7 GB | 6.7 GB | — |
MLX 2-bit | mlx | 2.25 | 8.0 GB | 8.0 GB | — |
Generic weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead. Runtime-specific artifacts use their published resident and download footprints; streamed models can therefore require much more disk than memory. Quality retention is an assumption, not a measured evaluation. Bonsai 2's bits per weight describe storage packing of the same ternary weights, not quality tiers. Its GGUF packs require Prism ML's llama.cpp fork, and its MLX pack requires the loader bundled with the model.
Coding & quality benchmarksCompare coding results →
| Benchmark | Category | Score | Run details | Reported by | Date | Source |
|---|---|---|---|---|---|---|
| BFCL v3 | agentic | 74.92 | thinking | publisher | — | link |
| HumanEval+ | coding | 95.12 | thinking | publisher | — | link |
| LiveCodeBench | coding | 90.07 | thinking | publisher | — | link |
| MBPP+ | coding | 83.07 | thinking | publisher | — | link |
| IFBench (prompt-loose) | general | 74 | thinking | publisher | — | link |
| IFEval | general | 91.31 | thinking | publisher | — | link |
| MMLU-Redux | general | 89.09 | thinking | publisher | — | link |
| AIME25 | reasoning | 95 | thinking | publisher | — | link |
| AIME26 | reasoning | 95.83 | thinking | publisher | — | link |
| GSM8K | reasoning | 96.66 | thinking | publisher | — | link |
| MATH-500 | reasoning | 98.8 | thinking | publisher | — | link |
| MuSR | reasoning | 70.63 | thinking | publisher | — | link |
| MMMU-Pro | vision | 75.49 | thinking | publisher | — | link |
| OCR Bench v2 | vision | 56.88 | thinking | publisher | — | link |
We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
91.1 t/s · $2,071
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
91.1 t/s · $2,071
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
91.1 t/s · $2,071
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
129.9 t/s measured at PQ2_0 · $3,332
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per purchase-price unit
91.1 t/s · $2,071
Highest decode tokens/second per 1,000 units of the displayed purchase currency. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Ternary Bonsai 2 27B2 measured, 0 estimated
2 of 2 rows
| Hardware↕ | Quant | Memory↕ | Decode▼ | Prefill↕ | Context | Price↕ | Fit | Confidence↕ |
|---|---|---|---|---|---|---|---|---|
| RTX 5090 workstation (1x 32 GB) NVIDIA · 32 GB · 1,792 GB/s | PTQ1_0recommended measured: PQ2_0 | 7.2 GB | 121 t/s | 1810 t/s | 256K | $3,332 | Comfortable | Low |
| Used RTX 4090 workstation (24 GB) NVIDIA · 24 GB · 1,008 GB/s | PTQ1_0recommended measured: PQ2_0 | 7.2 GB | 91.1 t/s | 1650 t/s | 192K | $2,071 | Comfortable | Low |
Similar models