Coding benchmarks

Published results for models in this catalogue. These measure a model together with an agent, prompt, tool setup, and compute budget—not the bare checkpoint alone.

Published results
9
Benchmarks represented
5
Models with results
4
Do not read this as one leaderboard. Compare scores only when the benchmark version, metric, harness, and test-time budget match. SWE-bench Verified is retained for historical context but is not a current frontier ranking signal.
Published coding and agentic results
ModelBenchmarkScoreRun detailsSourceDate
GLM-5.3DeepSWE66.9Not supplied by sourcesource28 Aug 2026
GLM-5.3 FlashDeepSWE63.4Not supplied by sourcesource26 Aug 2026
MiniMax M3SWE-bench Pro59Not supplied by sourcesource23 Jun 2026
Kimi K2.6SWE-bench Pro58.6Not supplied by sourcesource20 Apr 2026
MiniMax M3SWE-bench Verified · legacy80.5Not supplied by sourcesource23 Jun 2026
Kimi K2.6SWE-bench Verified · legacy80.2Not supplied by sourcesource20 Apr 2026
Kimi K2.6Terminal-Bench 2.066.7Not supplied by sourcesource20 Apr 2026
GLM-5.3Terminal-Bench 2.188.2Not supplied by sourcesource31 Aug 2026
GLM-5.3 FlashTerminal-Bench 2.184.3Not supplied by sourcesource26 Aug 2026
Scores are imported from canonical model repositories when available. A blank run-details field means the publisher did not provide structured configuration metadata; it is not assumed.