DeepSWE 1.1
Progress Over Time
Interactive timeline showing model performance evolution on DeepSWE 1.1
State-of-the-art frontier
Open
Proprietary
DeepSWE 1.1 Leaderboard
19 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 2 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 2 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 4 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 5 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 6 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 6 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 8 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 9 | Qwen3.8 MaxNew Alibaba Cloud / Qwen Team | 2.4T | — | — | ||
| 10 | xAI | — | 500K | $2.00 / $6.00 | ||
| 10 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 12 | Meta | — | 1.0M | $1.25 / $4.25 | ||
| 13 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 14 | Google | — | 1.0M | $1.50 / $7.50 | ||
| 15 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 16 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 17 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 | ||
| 18 | Anthropic | — | 200K | $3.00 / $15.00 | ||
| 19 | Google | — | 1.0M | $2.50 / $15.00 |
Notice missing or incorrect data?
What is DeepSWE 1.1?
DeepSWE 1.1 evaluates software engineering agents using the mini-swe-agent harness.
DeepSWE 1.1 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 19 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
GPT-5.6 Sol from OpenAI currently leads the DeepSWE 1.1 leaderboard with a score of 0.730 across 19 evaluated AI models.
FAQ
Common questions about the DeepSWE 1.1 benchmark and leaderboard.