The AI arena is free today

Open Playground

DeepSWE 1.1

Progress Over Time

Interactive timeline showing model performance evolution on DeepSWE 1.1

State-of-the-art frontier
Open
Proprietary

DeepSWE 1.1 Leaderboard

19 models
ContextCostLicense
11.1M$5.00 / $30.00
21.1M$2.00 / $12.00
21.0M$10.00 / $50.00
4
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
5
Anthropic
Anthropic
1.0M$5.00 / $25.00
61.1M$0.20 / $1.20
6
OpenAI
OpenAI
1.1M$5.00 / $30.00
81.0M$5.00 / $25.00
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T
10500K$2.00 / $6.00
101.0M$2.00 / $10.00
121.0M$1.25 / $4.25
13
OpenAI
OpenAI
1.0M$2.50 / $15.00
141.0M$1.50 / $7.50
15
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
161.0M$1.50 / $9.00
17
Moonshot AI
Moonshot AI
1.0T262K$0.74 / $3.50
18200K$3.00 / $15.00
191.0M$2.50 / $15.00
Notice missing or incorrect data?
About this benchmark

What is DeepSWE 1.1?

DeepSWE 1.1 evaluates software engineering agents using the mini-swe-agent harness.

DeepSWE 1.1 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 19 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the DeepSWE 1.1 leaderboard with a score of 0.730 across 19 evaluated AI models.

1GPT-5.6 SolOpenAI73.0%
2GPT-5.6 TerraOpenAI70.0%
2Claude Fable 5Anthropic70.0%
OSSKimi K3#4 open-weight69.0%

FAQ

Common questions about the DeepSWE 1.1 benchmark and leaderboard.

What is the DeepSWE 1.1 benchmark?

DeepSWE 1.1 evaluates software engineering agents using the mini-swe-agent harness.

What is the DeepSWE 1.1 leaderboard?

The DeepSWE 1.1 leaderboard ranks 19 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.730. The average score across all models is 0.535.

What is the highest DeepSWE 1.1 score?

The highest DeepSWE 1.1 score is 0.730, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on DeepSWE 1.1?

19 models have been evaluated on the DeepSWE 1.1 benchmark, with 0 verified results and 2 self-reported results.

What categories does DeepSWE 1.1 cover?

DeepSWE 1.1 is categorized under agents and code. The benchmark evaluates text models.

What's the difference between DeepSWE 1.1 and DeepSWE?

DeepSWE 1.1 is a variant of DeepSWE. See the DeepSWE leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on DeepSWE 1.1?

Kimi K3 by Moonshot AI is the top-ranked open-source model on DeepSWE 1.1, with a score of 0.690 (rank #4).

Which model offers the best value on DeepSWE 1.1?

Among models scoring within 10% of the leader, GPT-5.6 Luna from OpenAI is the cheapest, at $0.20 per million input tokens with a score of 0.670.

How recent are the DeepSWE 1.1 leaderboard results?

The DeepSWE 1.1 leaderboard was last updated in August 2026 and currently includes 19 evaluated models.