FrontierSWE

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on FrontierSWE

State-of-the-art frontier
Open
Proprietary

FrontierSWE Leaderboard

14 models
ContextCostLicense
11.0M$10.00 / $50.00
2
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
31.0M$5.00 / $25.00
4
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
5
OpenAI
OpenAI
1.1M$5.00 / $30.00
61.0M$5.00 / $25.00
71.0M$5.00 / $25.00
8
OpenAI
OpenAI
1.0M$2.50 / $15.00
91.0M$2.50 / $15.00
10
Zhipu AI
Zhipu AI
754B200K$1.40 / $4.40
111.6T1.0M$1.60 / $3.20
12
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
13
Moonshot AI
Moonshot AI
1.0T
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
Notice missing or incorrect data?
About this benchmark

What is FrontierSWE?

FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.

FrontierSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 14 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.9.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Claude Fable 5 from Anthropic currently leads the FrontierSWE leaderboard with a score of 0.900 across 14 evaluated AI models.

1Claude Fable 5Anthropic90.0%
2Kimi K3Moonshot AI81.2%
3Claude Opus 4.8Anthropic75.0%

FAQ

Common questions about the FrontierSWE benchmark and leaderboard.

What is the FrontierSWE benchmark?

FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.

What is the FrontierSWE leaderboard?

The FrontierSWE leaderboard ranks 14 AI models based on their performance on this benchmark. Currently, Claude Fable 5 by Anthropic leads with a score of 0.900. The average score across all models is 0.529.

What is the highest FrontierSWE score?

The highest FrontierSWE score is 0.900, achieved by Claude Fable 5 from Anthropic.

How many models are evaluated on FrontierSWE?

14 models have been evaluated on the FrontierSWE benchmark, with 0 verified results and 1 self-reported results.

Where can I find the FrontierSWE dataset?

The FrontierSWE dataset is available at https://www.frontierswe.com/.

What categories does FrontierSWE cover?

FrontierSWE is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on FrontierSWE?

Kimi K3 by Moonshot AI is the top-ranked open-source model on FrontierSWE, with a score of 0.812 (rank #2).

Which model offers the best value on FrontierSWE?

Among models scoring within 10% of the leader, Kimi K3 from Moonshot AI is the cheapest, at $3.00 per million input tokens with a score of 0.812.

How recent are the FrontierSWE leaderboard results?

The FrontierSWE leaderboard was last updated in July 2026 and currently includes 14 evaluated models.