Devstral Small 24B on RTX 3090 workstation (Ryzen 9 7950X, 64 GB)
23.572B on 24 GB at 936 GB/s. Hardware details · Model details
Yes. Devstral Small 24B fits on RTX 3090 workstation (Ryzen 9 7950X, 64 GB) at the recommended
Q4_K_M configuration, requiring approximately 16.4 GB at 8K context. Practical context capacity is 32K. Expected decode for the recommended configuration is ~38.8–55.8 t/s. EstimatedCompatibilityComfortable
Calculated fit
Yes
Calculated Uses under 80% of usable memory. Room for a long context and other work.
Recommended quantization
Q4_K_M
Calculated gguf
Memory required
16.4 GB
Calculated of 22.1 GB usable — 74%
Max practical context
32K
Calculated model supports 128K
Memory budget at 8K context
Model weights ⓘ
13.7 GB Calculated
KV cache ⓘ
1.3 GB Calculated
Runtime overhead ⓘ
1.4 GB Estimated
Total required
16.4 GB Calculated
Headroom ⓘ
5.7 GB Calculated
Highest-precision quantization that leaves headroom: uses 74% of usable memory at 8K context.
Discrete GPU: 24 GB of VRAM, of which we assume 92% is usable after driver and context overhead.
Performance
Estimated performance · recommended Q4_K_M
Decode, prefill and TTFT below are estimates for Q4_K_M. We do not have a comparable Q4_K_M measurement on this machine.
Estimated decode
~38.8–55.8 t/s
Estimated Q4_K_M; calculated range
Estimated prefill
~627–1300 t/s
Estimated Q4_K_M; calculated range
Estimated TTFT at 8K
~6.4–13.2 s
Estimated Q4_K_M; calculated range
Hardware load reference
600 W
Official spec hardware TDP; not this model run
How the estimate is calculated
- Decode: reading 23.6B active parameters at 4.85 bits/weight takes 20.36 ms at 936 GB/s x 75% achieved efficiency.
- Prefill: 142 TFLOPS (FP16) x 32% assumed model-FLOPs utilisation, divided by 2 x 23.6B parameters per token.
- Prefill throughput is highly engine-dependent. Flash attention, batch size and quantized KV all move this number substantially, and for sparse mixture-of-experts models it is the least reliable figure we produce.
- This is a calculated estimate, not a measurement. It assumes a single request, a short prompt, no speculative decoding and a warm model already resident in memory.
Estimator version estimator_v1. Stored with every estimated row so old estimates can be regenerated when the model improves.
Submit a benchmarkSubmissions stay pending until reviewed.
This model at other quantizations on this machine
| Quantization | Resident | Disk | Total needed | Utilisation | Fit | Max context |
|---|---|---|---|---|---|---|
Q4_K_Mrecommended | 13.7 GB | 13.7 GB | 16.4 GB | 74% | Comfortable | 32K |
Q5_K_M | 16.1 GB | 16.1 GB | 18.8 GB | 85% | Fits | 16K |
Q8_0 | 23.8 GB | 23.8 GB | 26.8 GB | 121% | Does not fit | 0 |
BF16 | 43.9 GB | 43.9 GB | 47.5 GB | 215% | Does not fit | 0 |
Cost of running Devstral Small 24B on this machineUnited States (federal) · C corporation · 8h/day
Monthly economic cost
$26
Calculated after tax
Codex/Claude Code
$100/mo
≈ $100/mo · local is $74 less
Net cash at purchase
$1,981
Calculated VAT not reclaimable
Total over 5 years
$1,578
Calculated after tax, after resale
Cost per USD/1M tokens
$4.39
Calculated 6.0M tokens/month
The $100 comparison uses the Codex Pro 5x / Claude Max 5x plans. A fixed planning conversion is used for USD. This compares monthly spend only: subscriptions have usage limits, local hardware has different capabilities and constraints, and taxes or regional pricing may change the charged amount. Prices checked 7 September 2026.
Purchase-price input: $1,981 in United States (federal), tax/VAT excluded · estimated · Used market exact-match reference (converted catalogue estimate). This is the localized purchase input; the after-tax economic cost is calculated separately below.
Monthly breakdown
| Depreciation 1,545 over 5 years, straight-line to a 436 residual | $25.75 |
| Electricity 31.3 kWh/month at 0.140/kWh | $4.38 |
| Cost of capital 4.0%/yr on 1,208 average capital employed | $4.03 |
| Monthly cost before tax | $34.16 |
| Electricity tax shield Running costs are deductible business expenses | −$0.92 |
| First-year expensing §179 (100.0%) | −$6.93 |
| Monthly economic cost after tax | $26.30 |
Three different numbers, deliberately
Cash cost
$1,981
Money that leaves the bank account on day one, net of reclaimable VAT.
Accounting depreciation
$25.75/month
$1,545 written down over 5 years to a $436 residual.
After-tax economic cost
$26.30/month
Depreciation plus running costs plus cost of capital, less the tax those deductions save. This is the figure to compare between machines.
Assumptions and sources (verified 2026-09-07)
VAT rate
0% Official spec
VAT recoverable
0% Official spec
Federal C-corporation estimate at the flat 21% rate. Pass-through entities should use the sole-proprietor estimate as a rougher proxy.
Effective deduction rate
21.00% Calculated
Headline marginal rate 21.00%.
Depreciation
5 years, straight-line Assumption
Residual value
$436 (22%) Assumption
Extrapolated.
Electricity
$0.140/kWh Assumption
Average power draw
178 W Calculated
Load 600 W for 20% of powered hours, idle 72 W for the rest — a machine that is on is not generating tokens the whole time.
Investment allowances
§179 100.0% Official spec
Worth $416 in total — first-year expensing that replaces later tax depreciation.
Cost of capital
$4.03/month Assumption
- VAT rate: No US federal VAT · verified 2026-09-07
- Marginal tax rate: IRS Publication 542 — Corporations · verified 2026-09-07
- VAT recoverable fraction: site assumption · verified 2026-09-07
- Useful life: IRS Publication 946 — How To Depreciate Property · verified 2026-09-07
- Electricity price: Site assumption · verified 2026-09-07
Assumes the machine generates tokens 20% of its 176 powered hours per month. Cost per token scales inversely with this number — halve the utilisation and the cost per token doubles.
Federal planning estimate only, not tax advice. State and local income tax, sales/use tax and incentives are excluded. It assumes 100% business use, a Section 179 election, enough business income to use it, and that the full annual limit remains available.
Efficiency metrics
Decode per $1,000 spent
23.9 t/s
Calculated purchase price only
Decode per $100/month
179.7 t/s
Calculated after-tax ownership cost
Tokens per joule
0.27
Calculated same as tokens/s per watt
USD per 1M output tokens
$4.39
Calculated at 20% utilisation
Local versus hosted APIs
| Hosted model | USD/1M output | Break-even | API at your volume | Verdict | Comparison type |
|---|---|---|---|---|---|
| Qwen3.8 Flash Alibaba Cloud | $0.42 | 17.1M/mo | $9 | API cheaper | Different model |
| DeepSeek-V4.1 Flash DeepSeek | $0.60 | 14.6M/mo | $11 | API cheaper | Different model |
| DeepSeek-V4 Pro DeepSeek | $1.98 | 3.6M/mo | $43 | Local cheaper | Different model |
| DeepSeek V4 Pro (Together) Together AI | $4.40 | 1.2M/mo | $127 | Local cheaper | Different model |
| GLM-5.3 Z.ai | $4.40 | 1.7M/mo | $93 | Local cheaper | Different model |
| Claude Haiku 4.5 Anthropic | $5.00 | 2.0M/mo | $78 | Local cheaper | Different model |
Different-model comparison: the hosted model is not the model you would run locally. Treat this as a workload-quality trade-off, not a direct economic equivalence — the frontier model may complete a task in fewer tokens, or complete tasks the local model cannot. Break-even is the monthly output volume at which API spend equals the $26/month economic cost of owning this machine, assuming 8 input tokens per output token. Hosted prices are published in USD. This table is the canonical US default.
Devstral Small 24B on other hardware
| Hardware | Decode | Price | Fit |
|---|---|---|---|
| NVIDIA DGX H200 (8x H200, 1,128 GB) | ~354–509 t/s | $359,430 * | Comfortable |
| Quad RTX 5090 workstation (4x 32 GB) | ~87–125 t/s | $12,611 * | Comfortable |
| Lenovo ThinkStation PX (4x RTX PRO 6000, 384 GB) | ~76.3–110 t/s | $58,553 * | Comfortable |
| RTX 5090 workstation (1x 32 GB) | ~65–93.5 t/s | $3,332 * | Comfortable |
| Dual RTX 5090 workstation (2x 32 GB) | ~58.5–84.2 t/s | $5,854 * | Comfortable |
| Dell Precision 7960 Rack (2x RTX PRO 6000, 192 GB) | ~53.7–77.3 t/s | $36,663 * | Comfortable |
Nearest alternatives to the RTX 3090 workstation (Ryzen 9 7950X, 64 GB)
Devstral Small 24B on Used RTX 4090 workstation (24 GB)24 GBDevstral Small 24B on RTX 5090 workstation (1x 32 GB)32 GBDevstral Small 24B on Dual used RTX 4090 workstation (48 GB)48 GBDevstral Small 24B on Radeon AI PRO R9700 workstation (32 GB)32 GBDevstral Small 24B on Dual RTX 5090 workstation (2x 32 GB)64 GB
Compare RTX 3090 workstation (Ryzen 9 7950X, 64 GB) against Used RTX 4090 workstation (24 GB) →
Compare RTX 3090 workstation (Ryzen 9 7950X, 64 GB) against Used RTX 4090 workstation (24 GB) →