Ternary Bonsai 2 27B on Dual RTX 5090 workstation (2x 32 GB)

27.36B on 64 GB at 3,584 GB/s. Hardware details · Model details

Run locally →
Yes. Ternary Bonsai 2 27B fits on Dual RTX 5090 workstation (2x 32 GB) at the recommended PTQ1_0 configuration, requiring approximately 7.6 GB at 8K context. Practical context capacity is 256K.
CompatibilityComfortable
Calculated fit
Yes
Calculated Uses under 80% of usable memory. Room for a long context and other work.
Recommended quantization
PTQ1_0
Calculated gguf
Memory required
7.6 GB
Calculated of 56.3 GB usable — 14%
Max practical context
256K
Calculated model supports 256K

Memory budget at 8K context

Model weights
5.5 GB Measured
KV cache
0.5 GB Calculated
Runtime overhead
1.6 GB Estimated
Total required
7.6 GB Calculated
Headroom
48.7 GB Calculated
Smallest published ternary pack that leaves headroom: uses 14% of usable memory at 8K context.
2 GPUs providing 64 GB aggregate VRAM. Assumes tensor- or layer-parallel sharding; each GPU carries its own context and communication buffers.
Hybrid attention: token-growing KV memory is charged to 16 full-attention blocks. The linear-attention blocks also hold fixed-size state, so the memory estimate is approximate.
Multi-GPU throughput depends heavily on interconnect. Without NVLink or NVSwitch, tensor parallelism over PCIe adds latency that partially offsets the extra bandwidth.
Performance

Estimated performance · recommended PTQ1_0

Decode, prefill and TTFT below are estimates for PTQ1_0. We do not have a comparable PTQ1_0 measurement on this machine.

Estimated decode
Estimated PTQ1_0; calculated range
Estimated prefill
Estimated PTQ1_0; calculated range
Estimated TTFT at 8K
~1.1–2.4 s
Estimated PTQ1_0; calculated range
Hardware load reference
1180 W
Measured machine-level load; not this model run
No throughput figures for this configuration. We do not estimate performance for a model that cannot load.
How the estimate is calculated
  1. Multi-GPU: 1792 GB/s per card, with each additional card contributing 40% of its bandwidth — 2509 GB/s effective, not the 3584 GB/s aggregate. Cross-GPU collectives use PCIe rather than a dedicated GPU fabric.
  2. Decode: reading 27.4B active parameters at 1.75 bits/weight takes 3.01 ms at 2509 GB/s x 79% achieved efficiency.
  3. Per-token overhead of 1.4 ms (kernel launches, attention bookkeeping, sampling) is significant here — this model is not purely bandwidth-bound on this hardware.
  4. Prefill: 419 TFLOPS (FP16) x 4 for native FP4 tensor cores x 0.71 calibrated against measured prefill on this platform x 26% assumed model-FLOPs utilisation, divided by 2 x 27.4B parameters per token.

  • With few active parameters on fast memory, fixed per-token overhead rather than bandwidth sets the ceiling. Real engines vary widely in how well they hide it.
  • Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
  • This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.

Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.

Submit a benchmarkSubmissions stay pending until reviewed.
This model at other quantizations on this machine
QuantizationResidentDiskTotal neededUtilisationFitMax context
PTQ1_0recommended5.5 GB5.5 GB7.6 GB14%Comfortable256K
PQ2_06.7 GB6.7 GB8.8 GB16%Comfortable256K
Cost of running Ternary Bonsai 2 27B on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$80
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $20 less
Net cash at purchase
$5,854
Calculated VAT not reclaimable
Total over 5 years
$4,799
Calculated after tax, after resale
Cost per USD/1M tokens
Calculated

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Purchase-price input: $5,854 in United States (federal), tax/VAT excluded · estimated · NL retail (system build) (converted catalogue estimate). This is the localized purchase input; the after-tax economic cost is calculated separately below.

Monthly breakdown

Depreciation
4,976 over 5 years, straight-line to a 878 residual
$82.94
Electricity
57.0 kWh/month at 0.140/kWh
$7.98
Cost of capital
4.0%/yr on 3,366 average capital employed
$11.22
Monthly cost before tax$102.14
Electricity tax shield
Running costs are deductible business expenses
−$1.68
First-year expensing
§179 (100.0%)
−$20.49
Monthly economic cost after tax$79.98

Three different numbers, deliberately

Cash cost
$5,854
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$82.94/month
$4,976 written down over 5 years to a $878 residual.
After-tax economic cost
$79.98/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$878 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
324 W Calculated
Load 1180 W for 20% of powered hours, idle 110 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $1,229 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$11.22/month Assumption
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.