The AI arena is free today

Open Superagent

DeepSWE

Progress Over Time

Interactive timeline showing model performance evolution on DeepSWE

State-of-the-art frontier
Open
Proprietary

DeepSWE Leaderboard

12 models
ContextCostLicense
11.1M$5.00 / $30.00
21.1M$2.00 / $12.00
3
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
41.1M$0.20 / $1.20
51.6T1.0M$0.43 / $0.87
61.0M$0.22 / $0.66
7304B1.0M$0.09 / $0.18
8500K$2.00 / $6.00
9
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
10
ByteDance
ByteDance
11
Tencent
Tencent
295B
12
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is DeepSWE?

DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.

DeepSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727 across 12 evaluated AI models.

1GPT-5.6 SolOpenAI72.7%
2GPT-5.6 TerraOpenAI69.6%
3Kimi K3Moonshot AI67.5%

FAQ

Common questions about the DeepSWE benchmark and leaderboard.

What is the DeepSWE benchmark?

DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.

What is the DeepSWE leaderboard?

The DeepSWE leaderboard ranks 12 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.727. The average score across all models is 0.530.

What is the highest DeepSWE score?

The highest DeepSWE score is 0.727, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on DeepSWE?

12 models have been evaluated on the DeepSWE benchmark, with 0 verified results and 12 self-reported results.

What categories does DeepSWE cover?

DeepSWE is categorized under agents and code. The benchmark evaluates text models.

Are there variants of DeepSWE?

Yes. DeepSWE has 2 related variants: DeepSWE 1.0, DeepSWE 1.1.

What is the best open-source model on DeepSWE?

Kimi K3 by Moonshot AI is the top-ranked open-source model on DeepSWE, with a score of 0.675 (rank #3).

Which model offers the best value on DeepSWE?

Among models scoring within 10% of the leader, GPT-5.6 Luna from OpenAI is the cheapest, at $0.20 per million input tokens with a score of 0.672.

How recent are the DeepSWE leaderboard results?

The DeepSWE leaderboard was last updated in August 2026 and currently includes 12 evaluated models.