FrontierSWE
Progress Over Time
Interactive timeline showing model performance evolution on FrontierSWE
FrontierSWE Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 2 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 3 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 4 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 5 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 6 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 7 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 8 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 9 | Google | — | 1.0M | $2.50 / $15.00 | ||
| 10 | Zhipu AI | 754B | 200K | $1.40 / $4.40 | ||
| 11 | DeepSeek | 1.6T | 1.0M | $1.60 / $3.20 | ||
| 12 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 13 | Moonshot AI | 1.0T | — | — | ||
| 14 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 |
What is FrontierSWE?
FrontierSWE measures whether an agent can complete open-ended technical projects at the scale of hours to tens of hours, spanning systems optimization, large-scale code construction, and applied ML research. Performance is reported as a dominance score, where higher is better.
FrontierSWE is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 14 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.9.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Claude Fable 5 from Anthropic currently leads the FrontierSWE leaderboard with a score of 0.900 across 14 evaluated AI models.
FAQ
Common questions about the FrontierSWE benchmark and leaderboard.