DeepSeek-V4.1 Flash

DeepSeek's new multimodal Causal Encoder-Decoder MoE has a 552B backbone, activates 8B parameters for input and 16B for output, and supports a one-million-token context. The pinned archive contains about 763B tensors once its 196B Engram memory, vision stack, and auxiliary weights are counted, so even 4-bit local inference is a cluster-scale proposition. Its native FP4 cache is unusually compact at 890 bytes per token.

DeepSeekMixture of expertsagenticreasoning2 machines can run it
Total parameters
763.205B
Official spec sets memory need
Active parameters
16B
Official spec sets decode speed
Context
1M
Official spec tokens
4-bit weights
416 GB
Calculated before KV cache
This is a sparse mixture of experts. All 763.205B of weights must sit in memory, but only 16B are read per generated token — so it needs the memory of a 763.205B model and generates at roughly the speed of a 16B one. That is why large unified-memory machines suit it and fast 32 GB GPUs do not.
Model specification
Publisher
DeepSeek Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
763.205B Official spec
Active parameters
16B per token Official spec
Context length
1,048,576 tokens Official spec
Attention
CED · 1 KV heads × 512 dim × 40 layers Official spec
Licence
MIT Official spec source
Specification confidence
High · verified 2026-09-10 Official spec source
Released
10 Sept 2026 Official spec
Official source
Quantizations and memory
QuantizationFormatBits/weightWeightsQuality kept
MLX 4-bitmlx4.5415.8 GB98.0%
Q4_K_Mdefaultgguf4.85443.8 GB98.5%
Q8_0gguf8.5770.3 GB99.9%
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Coding & quality benchmarksCompare coding results →

No published benchmark results have been imported for this model yet. This is missing evidence, not a score of zero.

We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
~70–101 t/s · $90,082
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~70–101 t/s · $90,082
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~70–101 t/s · $90,082
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per purchase-price unit
~451–649 t/s · $359,430
Highest decode tokens/second per 1,000 units of the displayed purchase currency. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs DeepSeek-V4.1 Flash0 measured, 2 estimated
2 of 2 rows
HardwareQuantMemoryDecodePrefillContextPriceFitConfidence
NVIDIA DGX H200 (8x H200, 1,128 GB)
NVIDIA · 1,128 GB · 38,400 GB/s
Q4_K_M461.0 GB~451–649 t/s~24210–50290 t/s1024K$359,430ComfortableEstimated
NVIDIA DGX Station GB300 (748 GB)
NVIDIA · 748 GB · 7,100 GB/s
Q4_K_M458.2 GB~70–101 t/s~3240–6720 t/s1024K$90,082ComfortableEstimated
Hosted alternatives for this exact model

These endpoints serve the same open weights, so comparing them against local ownership is a like-for-like economic question rather than a quality trade-off.

ProviderInput /1MOutput /1MVerifiedSource
DeepSeek$0.15$0.6010 Sept 2026link
Similar models