The AI arena is free today

Open Superagent

SWE-Bench Pro

Progress Over Time

Interactive timeline showing model performance evolution on SWE-Bench Pro

State-of-the-art frontier
Open
Proprietary

SWE-Bench Pro Leaderboard

58 models
ContextCostLicense
11.0M$10.00 / $50.00
2
31.0M$5.00 / $25.00
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
5770B
6500K$2.00 / $6.00
71.1M$5.00 / $30.00
81.0M$5.00 / $25.00
91.1M$2.00 / $12.00
101.0M$2.00 / $10.00
111.1M$0.20 / $1.20
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
14
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.40 / $3.00
161.0M$1.25 / $4.25
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
18
Shanghai AI Laboratory
Shanghai AI Laboratory
744B
19
Poolside
Poolside
118B1.0M$0.10 / $0.20
20
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
211.0M$1.50 / $7.50
22
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
22
OpenAI
OpenAI
1.1M$5.00 / $30.00
24
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
25
Tencent
Tencent
295B262K$0.14 / $0.58
26
OpenAI
OpenAI
1.0M$2.50 / $15.00
27
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28
ByteDance
ByteDance
291.0T1.0M$0.43 / $0.87
30
31400K$1.75 / $14.00
32
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
32
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
34400K$1.75 / $14.00
35198B262K$0.20 / $1.15
36205K$0.30 / $1.20
37
Xiaomi
Xiaomi
311B1.0M$0.17 / $0.34
38
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
391.6T1.0M$1.30 / $2.60
39230B1.0M$0.30 / $1.20
411.0M$1.50 / $9.00
42400K$0.75 / $4.50
431.0M$2.00 / $12.00
431.0M$0.30 / $2.50
45
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
461.0T
47284B1.0M$0.09 / $0.18
48400K$0.20 / $1.25
48
50284B1.0M$0.09 / $0.18
150 of 58
1/2
Notice missing or incorrect data?
About this benchmark

What is SWE-Bench Pro?

SWE-Bench Pro is an advanced version of SWE-Bench that evaluates language models on complex, real-world software engineering tasks requiring extended reasoning and multi-step problem solving.

SWE-Bench Pro is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 58 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.

Current leaders

Claude Fable 5 from Anthropic currently leads the SWE-Bench Pro leaderboard with a score of 0.800 across 58 evaluated AI models.

1Claude Fable 5Anthropic80.0%
2Claude Mythos PreviewAnthropic77.8%
3Claude Opus 4.8Anthropic69.2%
OSSHy4 preview#5 open-weight65.7%

FAQ

Common questions about the SWE-Bench Pro benchmark and leaderboard.

What is the SWE-Bench Pro benchmark?

SWE-Bench Pro is an advanced version of SWE-Bench that evaluates language models on complex, real-world software engineering tasks requiring extended reasoning and multi-step problem solving.

What is the SWE-Bench Pro leaderboard?

The SWE-Bench Pro leaderboard ranks 58 AI models based on their performance on this benchmark. Currently, Claude Fable 5 by Anthropic leads with a score of 0.800. The average score across all models is 0.571.

What is the highest SWE-Bench Pro score?

The highest SWE-Bench Pro score is 0.800, achieved by Claude Fable 5 from Anthropic.

How many models are evaluated on SWE-Bench Pro?

58 models have been evaluated on the SWE-Bench Pro benchmark, with 0 verified results and 58 self-reported results.

What categories does SWE-Bench Pro cover?

SWE-Bench Pro is categorized under reasoning, agents, and code. The benchmark evaluates text models.

What's the difference between SWE-Bench Pro and SWE-Bench Verified?

SWE-Bench Pro is a variant of SWE-Bench Verified. See the SWE-Bench Verified leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on SWE-Bench Pro?

Hy4 preview by Tencent is the top-ranked open-source model on SWE-Bench Pro, with a score of 0.657 (rank #5).

Which model offers the best value on SWE-Bench Pro?

Among models scoring within 10% of the leader, Claude Fable 5 from Anthropic is the cheapest, at $10.00 per million input tokens with a score of 0.800.

How recent are the SWE-Bench Pro leaderboard results?

The SWE-Bench Pro leaderboard was last updated in September 2026 and currently includes 58 evaluated models.