DeepSWE
Progress Over Time
Interactive timeline showing model performance evolution on DeepSWE
DeepSWE Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 2 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 3 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 4 | OpenAI | — | 1.1M | $1.00 / $6.00 | ||
| 5 | xAI | — | 500K | $2.00 / $6.00 | ||
| 6 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 7 | ByteDance | — | — | — | ||
| 8 | Tencent | 295B | — | — | ||
| 9 | ByteDance | — | — | — |
Sub-benchmarks
What is DeepSWE?
DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.
DeepSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 9 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727 across 9 evaluated AI models.
FAQ
Common questions about the DeepSWE benchmark and leaderboard.