The AI arena is free today

Open Superagent

FrontierSWE

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on FrontierSWE

State-of-the-art frontier
Open
Proprietary

FrontierSWE Leaderboard

16 models
ContextCostLicense
11.0M$10.00 / $50.00
2
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
3
Zhipu AI
Zhipu AI
753B1.0M$1.40 / $4.40
41.0M$5.00 / $25.00
5
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
7
OpenAI
OpenAI
1.1M$5.00 / $30.00
81.0M$5.00 / $25.00
91.0M$5.00 / $25.00
10
OpenAI
OpenAI
1.0M$2.50 / $15.00
111.0M$2.50 / $15.00
12
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
131.6T
14
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
15
Moonshot AI
Moonshot AI
1.0T
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
Notice missing or incorrect data?
About this benchmark

What is FrontierSWE?

FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.

FrontierSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 16 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.9.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Claude Fable 5 from Anthropic currently leads the FrontierSWE leaderboard with a score of 0.900 across 16 evaluated AI models.

1Claude Fable 5Anthropic90.0%
2Kimi K3Moonshot AI81.2%
3GLM-5.3Zhipu AI78.1%

FAQ

Common questions about the FrontierSWE benchmark and leaderboard.

What is the FrontierSWE benchmark?

FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.

What is the FrontierSWE leaderboard?

The FrontierSWE leaderboard ranks 16 AI models based on their performance on this benchmark. Currently, Claude Fable 5 by Anthropic leads with a score of 0.900. The average score across all models is 0.558.

What is the highest FrontierSWE score?

The highest FrontierSWE score is 0.900, achieved by Claude Fable 5 from Anthropic.

How many models are evaluated on FrontierSWE?

16 models have been evaluated on the FrontierSWE benchmark, with 0 verified results and 3 self-reported results.

Where can I find the FrontierSWE dataset?

The FrontierSWE dataset is available at https://www.frontierswe.com/.

What categories does FrontierSWE cover?

FrontierSWE is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on FrontierSWE?

Kimi K3 by Moonshot AI is the top-ranked open-source model on FrontierSWE, with a score of 0.812 (rank #2).

Which model offers the best value on FrontierSWE?

Among models scoring within 10% of the leader, Kimi K3 from Moonshot AI is the cheapest, at $3.00 per million input tokens with a score of 0.812.

How recent are the FrontierSWE leaderboard results?

The FrontierSWE leaderboard was last updated in August 2026 and currently includes 16 evaluated models.