DeepSWE

Progress Over Time

Interactive timeline showing model performance evolution on DeepSWE

State-of-the-art frontier
Open
Proprietary

DeepSWE Leaderboard

9 models
ContextCostLicense
11.1M$5.00 / $30.00
21.1M$2.50 / $15.00
3
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
41.1M$1.00 / $6.00
5500K$2.00 / $6.00
6
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
7
ByteDance
ByteDance
8
Tencent
Tencent
295B
9
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is DeepSWE?

DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.

DeepSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 9 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the DeepSWE leaderboard with a score of 0.727 across 9 evaluated AI models.

1GPT-5.6 SolOpenAI72.7%
2GPT-5.6 TerraOpenAI69.6%
3Kimi K3Moonshot AI67.5%

FAQ

Common questions about the DeepSWE benchmark and leaderboard.

What is the DeepSWE benchmark?

DeepSWE is a software engineering agent benchmark evaluated with the mini-swe-agent harness, where each task is solved in an isolated container with no internet access. It measures an agent's ability to autonomously resolve real-world coding issues end to end.

What is the DeepSWE leaderboard?

The DeepSWE leaderboard ranks 9 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.727. The average score across all models is 0.511.

What is the highest DeepSWE score?

The highest DeepSWE score is 0.727, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on DeepSWE?

9 models have been evaluated on the DeepSWE benchmark, with 0 verified results and 9 self-reported results.

What categories does DeepSWE cover?

DeepSWE is categorized under agents and code. The benchmark evaluates text models.

Are there variants of DeepSWE?

Yes. DeepSWE has 1 related variant: DeepSWE 1.0.

What is the best open-source model on DeepSWE?

Kimi K3 by Moonshot AI is the top-ranked open-source model on DeepSWE, with a score of 0.675 (rank #3).

Which model offers the best value on DeepSWE?

Among models scoring within 10% of the leader, GPT-5.6 Luna from OpenAI is the cheapest, at $1.00 per million input tokens with a score of 0.672.

How recent are the DeepSWE leaderboard results?

The DeepSWE leaderboard was last updated in July 2026 and currently includes 9 evaluated models.