Step 3.7 Flash 198B-A11B on RTX PRO 6000 Max-Q workstation (96 GB, 300 W)
201.365B (11B active) on 96 GB at 1,792 GB/s. Hardware details · Model details
No calculated fit. Step 3.7 Flash 198B-A11B does not fit under standard overhead assumptions at
Q4_K_M: the model calculates 132.9 GB required against 88.3 GB usable.CompatibilityDoes not fit
Calculated fit
No
Calculated Requires more than 97% of usable memory. Needs a smaller quantization or more memory.
Smallest catalogued quantization
Q4_K_M
Calculated gguf
Memory required
132.9 GB
Calculated of 88.3 GB usable — 150%
Max practical context
0
Calculated model supports 256K
Memory budget at 8K context
Model weights ⓘ
117.1 GB Calculated
KV cache ⓘ
11.3 GB Calculated
Runtime overhead ⓘ
4.5 GB Estimated
Total required
132.9 GB Calculated
Headroom ⓘ
-44.5 GB Calculated
No quantization in our catalogue fits this machine.
Discrete GPU: 96 GB of VRAM, of which we assume 92% is usable after driver and context overhead.
Mixture of experts: all 201.365B parameters must be resident in memory even though only ~11B are active per token. Memory follows total parameters; speed follows active parameters.
This model can be run with layers offloaded to system RAM, but decode throughput typically drops by an order of magnitude once any significant fraction of the weights crosses PCIe. We do not count offload as fitting.
Performance
Estimated performance · recommended Q4_K_M
Decode, prefill and TTFT below are estimates for Q4_K_M. We do not have a comparable Q4_K_M measurement on this machine.
Estimated decode
—
Estimated Q4_K_M; calculated range
Estimated prefill
—
Estimated Q4_K_M; calculated range
Estimated TTFT at 8K
~7.0–14.6 s
Estimated Q4_K_M; calculated range
Hardware load reference
390 W
Measured machine-level load; not this model run
No throughput figures for this configuration. We do not estimate performance for a model that cannot load.
How the estimate is calculated
- Decode: reading 11.0B active parameters at 4.85 bits/weight takes 5.01 ms at 1792 GB/s x 74% achieved efficiency.
- MoE routing penalty of 15% applied: expert gathers are less bandwidth-efficient than a dense sweep.
- Prefill: 190 TFLOPS (FP16) x 2 for native FP8 tensor cores x 0.67 calibrated against measured prefill on this platform x 32% assumed model-FLOPs utilisation, divided by 2 x 47.1B parameters per token.
- Prefill uses 47.1B effective parameters, not the 11B active in decode: a batch of hundreds of tokens routes across most of the expert pool.
- MoE decode depends on how well the engine batches expert gathers; real results vary more than for dense models.
- Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
- This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.
Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.
Submit a benchmarkSubmissions stay pending until reviewed.
This model at other quantizations on this machine
| Quantization | Resident | Disk | Total needed | Utilisation | Fit | Max context |
|---|---|---|---|---|---|---|
Q4_K_Msmallest listed | 117.1 GB | 117.1 GB | 132.9 GB | 150% | Does not fit | 0 |
Q5_K_M | 137.4 GB | 137.4 GB | 153.8 GB | 174% | Does not fit | 0 |
Q8_0 | 203.2 GB | 203.2 GB | 221.6 GB | 251% | Does not fit | 0 |
BF16 | 375.1 GB | 375.1 GB | 398.6 GB | 451% | Does not fit | 0 |
Cost of running Step 3.7 Flash 198B-A11B on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$138
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $38 more
Net cash at purchase
$10,809
Calculated VAT not reclaimable
Total over 5 years
$8,294
Calculated after tax, after resale
Cost per USD/1M tokens
—
Calculated
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Purchase-price input: $10,809 in United States (federal), tax/VAT excluded · estimated · NL workstation integrator (converted catalogue estimate). This is the localized purchase input; the after-tax economic cost is calculated separately below.
Monthly breakdown
| Depreciation 9,188 over 5 years, straight-line to a 1,621 residual | $153.13 |
| Electricity 20.1 kWh/month at 0.140/kWh | $2.81 |
| Cost of capital 4.0%/yr on 6,215 average capital employed | $20.72 |
| Monthly cost before tax | $176.65 |
| Electricity tax shield Running costs are deductible business expenses | −$0.59 |
| First-year expensing §179 (100.0%) | −$37.83 |
| Monthly economic cost after tax | $138.23 |
Three different numbers, deliberately
Cash cost
$10,809
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$153.13/month
$9,188 written down over 5 years to a $1,621 residual.
After-tax economic cost
$138.23/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$1,621 (15%) Assumption
Two GPU generations later, the card is worth a fraction of its list price.
Electricity
$0.140/kWh Assumption
Average power draw
114 W Calculated
Load 390 W for 20% of powered hours, idle 45 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $2,270 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$20.72/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.
Step 3.7 Flash 198B-A11B on other hardware
| Hardware | Decode | Price | Fit |
|---|---|---|---|
| NVIDIA DGX H200 (8x H200, 1,128 GB) | ~431–620 t/s | $359,430 * | Comfortable |
| Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB) | ~143–205 t/s | $36,663 * | Comfortable |
| Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB) | ~126–181 t/s | $58,553 * | Comfortable |
| Mac Studio M5 Ultra 256 GB | ~84.1–121 t/s | $8,647 * | Comfortable |
| NVIDIA DGX Station GB300 (748 GB) | ~58.9–84.8 t/s | $90,082 * | Comfortable |
| Mac Studio M5 Ultra 512 GB | ~58.3–83.8 t/s | $11,350 * | Comfortable |
Nearest alternatives to the RTX PRO 6000 Max-Q workstation (96 GB, 300 W)
Step 3.7 Flash 198B-A11B on RTX PRO 6000 Blackwell workstation (96 GB)96 GBStep 3.7 Flash 198B-A11B on Quad RTX 5090 workstation (4x 32 GB)128 GBStep 3.7 Flash 198B-A11B on Dual RTX 5090 workstation (2x 32 GB)64 GBStep 3.7 Flash 198B-A11B on Dual used RTX 4090 workstation (48 GB)48 GBStep 3.7 Flash 198B-A11B on RTX 5090 workstation (1x 32 GB)32 GB
Compare RTX PRO 6000 Max-Q workstation (96 GB, 300 W) against RTX PRO 6000 Blackwell workstation (96 GB) →
Compare RTX PRO 6000 Max-Q workstation (96 GB, 300 W) against RTX PRO 6000 Blackwell workstation (96 GB) →