Coding benchmarks
Published results for models in this catalogue. These measure a model together with an agent, prompt, tool setup, and compute budget—not the bare checkpoint alone.
Published results
9Benchmarks represented
5Models with results
4Do not read this as one leaderboard. Compare scores only when the benchmark version, metric, harness, and test-time budget match. SWE-bench Verified is retained for historical context but is not a current frontier ranking signal.
Published coding and agentic results
| Model | Benchmark | Score | Run details | Source | Date |
|---|---|---|---|---|---|
| GLM-5.3 | DeepSWE | 66.9 | Not supplied by source | source | 28 Aug 2026 |
| GLM-5.3 Flash | DeepSWE | 63.4 | Not supplied by source | source | 26 Aug 2026 |
| MiniMax M3 | SWE-bench Pro | 59 | Not supplied by source | source | 23 Jun 2026 |
| Kimi K2.6 | SWE-bench Pro | 58.6 | Not supplied by source | source | 20 Apr 2026 |
| MiniMax M3 | SWE-bench Verified · legacy | 80.5 | Not supplied by source | source | 23 Jun 2026 |
| Kimi K2.6 | SWE-bench Verified · legacy | 80.2 | Not supplied by source | source | 20 Apr 2026 |
| Kimi K2.6 | Terminal-Bench 2.0 | 66.7 | Not supplied by source | source | 20 Apr 2026 |
| GLM-5.3 | Terminal-Bench 2.1 | 88.2 | Not supplied by source | source | 31 Aug 2026 |
| GLM-5.3 Flash | Terminal-Bench 2.1 | 84.3 | Not supplied by source | source | 26 Aug 2026 |
Scores are imported from canonical model repositories when available. A blank run-details field means the publisher did not provide structured configuration metadata; it is not assumed.