DeepSWE
Progress Over Time
Interactive timeline showing model performance evolution on DeepSWE
DeepSWE Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 2 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 3 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 4 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 5 | DeepSeek | 1.6T | 1.0M | $0.43 / $0.87 | ||
| 6 | DeepSeek | — | 1.0M | $0.22 / $0.66 | ||
| 7 | DeepSeek | 304B | 1.0M | $0.09 / $0.18 | ||
| 8 | xAI | — | 500K | $2.00 / $6.00 | ||
| 9 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 10 | ByteDance | — | — | — | ||
| 11 | Tencent | 295B | — | — | ||
| 12 | ByteDance | — | — | — |
Sub-benchmarks
What is DeepSWE?
DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.
DeepSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727 across 12 evaluated AI models.
FAQ
Common questions about the DeepSWE benchmark and leaderboard.