GLM-5.3 Flash

A 320B sparse MoE with only 18B active parameters and a one-million-token context, released under MIT. The active parameter count is what makes it interesting locally: decode speed tracks 18B, while memory demands track 320B. That combination suits large unified-memory machines and rules out every 32 GB GPU.

Z.aiMixture of expertscodingreasoning4 machines can run it
Total parameters
321.323B
Official spec sets memory need
Active parameters
18B
Official spec sets decode speed
Context
1M
Official spec tokens
4-bit weights
175 GB
Calculated before KV cache
This is a sparse mixture of experts. All 321.323B of weights must sit in memory, but only 18B are read per generated token — so it needs the memory of a 321.323B model and generates at roughly the speed of a 18B one. That is why large unified-memory machines suit it and fast 32 GB GPUs do not.
Model specification
Publisher
Z.ai Official spec
Architecture
Sparse mixture of experts Official spec
Total parameters
321.323B Official spec
Active parameters
18B per token Official spec
Context length
1,048,576 tokens Official spec
Attention
MLA · 64 KV heads × 64 dim × 45 layers Official spec
Licence
MIT Official spec source
Specification confidence
High · verified 2026-09-07 Official spec source
Released
1 Jul 2026 Official spec
Official source
Quantizations and memory
QuantizationFormatBits/weightWeightsQuality kept
MLX 4-bitmlx4.5175.1 GB98.0%
Q4_K_Mdefaultgguf4.85186.9 GB98.5%
Q8_0gguf8.5324.3 GB99.9%
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Quality benchmarks
BenchmarkCategoryScoreReported byDateSource
Terminal-Bench 2.1agentic84.3canonical model_repository26 Aug 2026link
DeepSWEcoding63.4canonical model_repository26 Aug 2026link
We show source metrics rather than deriving one opaque quality number. Different benchmarks measure genuinely different things, and collapsing them into a single score would hide exactly the disagreements worth seeing.
Recommendations
Cheapest that can run it
~35.4–51 t/s · €6.299
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~35.4–51 t/s · €6.299
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~35.4–51 t/s · €6.299
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per euro
~62.1–89.3 t/s · €9.599
Highest decode tokens/second per EUR 1,000 of purchase price. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs GLM-5.3 Flash0 measured, 4 estimated
4 of 4 rows
HardwareQuantMemoryDecodePrefillContextPriceFitConfidence
Mac Studio M5 Ultra 256 GB
Apple · 256 GB · 1,200 GB/s
Q4_K_M193.9 GB~62.1–89.3 t/s~382–793 t/s32K€9.599FitsEstimated
Mac Studio M5 Ultra 512 GB
Apple · 512 GB · 1,200 GB/s
Q8_0335.4 GB~36.6–52.7 t/s~382–793 t/s1024K€12.599ComfortableEstimated
Mac Studio M3 Ultra 256 GB
Apple · 256 GB · 819 GB/s
Q4_K_M193.9 GB~35.4–51 t/s~92–191 t/s32K€6.299FitsEstimated
Mac Studio M3 Ultra 512 GB
Apple · 512 GB · 819 GB/s
Q8_0335.4 GB~20.6–29.7 t/s~92–191 t/s1024K€9.199ComfortableEstimated
Similar models