Qwen3.8 Flash Next on Dual used RTX 4090 workstation (48 GB)

180B (6B active) on 48 GB at 2,016 GB/s. Hardware details · Model details

CompatibilityDoes not fit
Fits
No
Calculated Requires more than 97% of usable memory. Needs a smaller quantization or more memory.
Recommended quantization
MLX 4-bit
Calculated mlx
Memory required
103.2 GB
Calculated of 42.2 GB usable — 244%
Max practical context
0
Calculated model supports 256K

Memory budget at 8K context

Model weights
98.1 GB Calculated
KV cache
0.8 GB Calculated
Runtime overhead
4.3 GB Estimated
Total required
103.2 GB Calculated
Usable memory
42.2 GB Assumption
Headroom
-60.9 GB Calculated
No quantization in our catalogue fits this machine.
2 GPUs pooling 48 GB total VRAM. Assumes tensor- or layer-parallel sharding; each GPU carries its own context and communication buffers.
Mixture of experts: all 180B parameters must be resident in memory even though only ~6B are active per token. Memory follows total parameters; speed follows active parameters.
Multi-GPU throughput depends heavily on interconnect. Without NVLink, tensor parallelism over PCIe adds latency that partially offsets the extra bandwidth.
PerformanceEstimated0/10
Decode (generation)
Estimated calculated, not measured
Prefill (prompt)
Estimated
TTFT at 8K
~1.2–2.5 s
Estimated time to first token
Power while generating
980 W
Measured
No throughput figures for this pairing. We do not estimate performance for a model that cannot load: a tokens-per-second number for a configuration that will never run is noise, not data. The alternatives below are machines that can actually run Qwen3.8 Flash Next.
How the estimate is calculated
  1. Multi-GPU: 1008 GB/s per card, with each additional card contributing 40% of its bandwidth — 1411 GB/s effective, not the 2016 GB/s aggregate. Consumer Blackwell has no NVLink, so every layer's all-reduce crosses PCIe.
  2. Decode: reading 6.0B active parameters at 4.50 bits/weight takes 3.07 ms at 1411 GB/s x 78% achieved efficiency.
  3. MoE routing penalty of 15% applied: expert gathers are less bandwidth-efficient than a dense sweep.
  4. Per-token overhead of 1.4 ms (kernel launches, attention bookkeeping, sampling) is significant here — this model is not purely bandwidth-bound on this hardware.
  5. Prefill: 330 TFLOPS (FP16) x 4 for native FP4 tensor cores x 26% assumed model-FLOPs utilisation, divided by 2 x 32.9B parameters per token.
  6. Prefill uses 32.9B effective parameters, not the 6B active in decode: a batch of hundreds of tokens routes across most of the expert pool.

  • MoE decode depends on how well the engine batches expert gathers; real results vary more than for dense models.
  • With few active parameters on fast memory, fixed per-token overhead rather than bandwidth sets the ceiling. Real engines vary widely in how well they hide it.
  • Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
  • This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.

Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.

This model at other quantizations on this machine
QuantizationWeightsTotal neededUtilisationFitMax context
MLX 4-bitrecommended98.1 GB103.2 GB244%Does not fit0
Q4_K_M104.7 GB110.0 GB260%Does not fit0
Q8_0181.7 GB189.3 GB448%Does not fit0
BF16335.3 GB347.5 GB823%Does not fit0
Cost of running Qwen3.8 Flash Next on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$48
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $52 less
Net cash at purchase
$3,692
Calculated VAT not reclaimable
Total over 5 years
$2,873
Calculated after tax, after resale
Cost per USD/1M tokens
Calculated

The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.

Monthly breakdown

Depreciation
2,880 over 5 years, straight-line to a 812 residual
$48.00
Electricity
47.9 kWh/month at 0.140/kWh
$6.70
Cost of capital
4.0%/yr on 2,252 average capital employed
$7.51
Monthly cost before tax$62.21
Electricity tax shield
Running costs are deductible business expenses
−$1.41
First-year expensing
§179 (100.0%)
−$12.92
Monthly economic cost after tax$47.88

Three different numbers, deliberately

Cash cost
$3,692
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$48.00/month
$2,880 written down over 5 years to a $812 residual.
After-tax economic cost
$47.88/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$812 (22%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
272 W Calculated
Load 980 W for 20% of powered hours, idle 95 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $775 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$7.51/month Assumption
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.
Page quality score (why this page is or is not indexed)

6/13 marked noindex. Generated pages are gated so we do not ask a search engine to rank a page with nothing computed to say. The directive is emitted in the page head via the metadata API, not in the body, so it is authoritative.

  • Has a memory-fit calculation (+3)
  • Has a price (+2)
  • Has an ownership economics calculation (+2)
  • Well connected (18 internal links) (+1)
  • No benchmark or estimate
  • Model does not fit and there is no measurement to explain why it is worth knowing (-2)
  • Fails a hard requirement: an indexable page needs both a compatibility calculation and at least one benchmark or estimate.