Llama 3.3 70B

The reference dense 70B. Its performance on every machine here is almost pure memory bandwidth divided by 40 GB, which makes it the cleanest possible test of a machine's decode capability. Superseded on quality, indispensable as a yardstick.

MetaDensegeneral16 machines can run it
Total parameters
70.554B
Official spec sets memory need
Active parameters
70.554B
Official spec sets decode speed
Context
128K
Official spec tokens
4-bit weights
38 GB
Calculated before KV cache
Model specification
Publisher
Meta Official spec
Architecture
Dense transformer Official spec
Total parameters
70.554B Official spec
Context length
131,072 tokens Official spec
Attention
GQA · 8 KV heads × 128 dim × 80 layers Official spec
Licence
Llama 3.3 Community License Official spec source
Specification confidence
High · verified 2026-09-07 Official spec source
Released
6 Dec 2024 Official spec
Official source
Quantizations and memory
QuantizationFormatBits/weightWeightsQuality kept
MLX 4-bitmlx4.538.4 GB98.0%
Q4_K_Mdefaultgguf4.8541.0 GB98.5%
Q8_0gguf8.571.2 GB99.9%
BF16safetensors16131.4 GB100.0%
Weight sizes are computed from the parameter count and bits per weight plus a format-specific overhead for the layers that stay at higher precision — not read from a specific published file. Quality retention is an assumption, not a measured evaluation.
Recommendations
Cheapest that can run it
~0.5–0.7 t/s · €1.699
Lowest purchase price among configurations where the model fits at some quantization in our catalogue. Speed is not considered.
Cheapest above 20 t/s
~38.4–55.2 t/s · €6.499
20 tokens/second is roughly the point at which generation keeps pace with reading. Below it, interactive use feels like waiting.
Cheapest above 40 t/s
~38.4–55.2 t/s · €6.499
40 tokens/second is the threshold most people describe as comfortable for coding agents, where output arrives faster than you can review it.
Fastest with real measurements
Nothing in the database qualifies.
Highest throughput among configurations with an actual published measurement rather than our estimate.
Best throughput per euro
~38.4–55.2 t/s · €6.499
Highest decode tokens/second per EUR 1,000 of purchase price. Ignores running costs and resale — see the economics section for the full picture.
Hardware that runs Llama 3.3 70B0 measured, 16 estimated
16 of 16 rows
HardwareQuantMemoryDecodePrefillContextPriceFitConfidence
Dual RTX 5090 workstation (2x 32 GB)
NVIDIA · 64 GB · 3,584 GB/s
MLX 4-bit43.5 GB~38.4–55.2 t/s~1420–2950 t/s16K€6.499ComfortableEstimated
Quad RTX 5090 workstation (4x 32 GB)
NVIDIA · 128 GB · 7,168 GB/s
Q8_078.0 GB~32.3–46.4 t/s~1420–2950 t/s64K€13.999ComfortableEstimated
RTX PRO 6000 Blackwell workstation (96 GB)
NVIDIA · 96 GB · 1,792 GB/s
Q4_K_M45.8 GB~24.9–35.8 t/s~461–958 t/s96K€11.499ComfortableEstimated
RTX PRO 6000 Max-Q workstation (96 GB, 300 W)
NVIDIA · 96 GB · 1,792 GB/s
Q4_K_M45.8 GB~24.9–35.8 t/s~378–784 t/s96K€11.999ComfortableEstimated
Mac Studio M5 Ultra 96 GB
Apple · 96 GB · 1,200 GB/s
Q4_K_M45.8 GB~19.7–28.3 t/s~329–683 t/s96K€6.599ComfortableEstimated
Mac Studio M5 Ultra 512 GB
Apple · 512 GB · 1,200 GB/s
Q8_076.8 GB~11.3–16.3 t/s~411–854 t/s128K€12.599ComfortableEstimated
Mac Studio M5 Ultra 256 GB
Apple · 256 GB · 1,200 GB/s
Q8_076.8 GB~11.3–16.3 t/s~411–854 t/s128K€9.599ComfortableEstimated
Mac Studio M5 Max 64 GB
Apple · 64 GB · 614 GB/s
MLX 4-bit43.1 GB~11–15.8 t/s~206–427 t/s16K€3.899ComfortableEstimated
Mac Studio M3 Ultra 512 GB
Apple · 512 GB · 819 GB/s
Q8_076.8 GB~6.3–9.1 t/s~99.2–206 t/s128K€9.199ComfortableEstimated
Mac Studio M3 Ultra 256 GB
Apple · 256 GB · 819 GB/s
Q8_076.8 GB~6.3–9.1 t/s~99.2–206 t/s128K€6.299ComfortableEstimated
Mac Studio M5 Max 128 GB
Apple · 128 GB · 614 GB/s
Q8_076.8 GB~5.8–8.4 t/s~206–427 t/s64K€4.799ComfortableEstimated
Mac mini M4 Pro 64 GB
Apple · 64 GB · 273 GB/s
MLX 4-bit43.1 GB~3.9–5.7 t/s~17.2–35.8 t/s16K€2.499ComfortableEstimated
NVIDIA DGX Spark 128 GB
NVIDIA · 128 GB · 273 GB/s
Q8_076.8 GB~2–2.9 t/s~177–368 t/s64K€4.299ComfortableEstimated
Framework Desktop (Ryzen AI Max+ 395, 128 GB)
AMD · 128 GB · 256 GB/s
Q8_076.8 GB~1.3–1.9 t/s~48.9–102 t/s64K€2.399ComfortableEstimated
GMKtec EVO-X2 (Ryzen AI Max+ 395, 128 GB)
AMD · 128 GB · 256 GB/s
Q8_076.8 GB~1.2–1.7 t/s~48.9–102 t/s64K€1.899ComfortableEstimated
CPU-only workstation (Ryzen 9950X, 192 GB DDR5)
Generic · 192 GB · 90 GB/s
Q8_076.8 GB~0.5–0.7 t/s~5.9–12.2 t/s128K€1.699ComfortableEstimated
Similar models