FrontierSWE V2
Progress Over Time
Interactive timeline showing model performance evolution on FrontierSWE V2
FrontierSWE V2 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $4.00 / $20.00 | ||
| 2 | Anthropic | — | 1.0M | $2.00 / $10.00 | ||
| 3 | Google | — | — | — | ||
| 4 | Anthropic | — | 1.0M | $0.10 / $0.50 | ||
| 5 | xAI | — | 500K | $2.00 / $6.00 |
What is FrontierSWE V2?
FrontierSWE V2 evaluates agents on open-ended technical projects and reports mean@5. Kept separate from FrontierSWE V1 dominance-score results, which are not comparable.
FrontierSWE V2 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Claude Opus 5.5 from Anthropic currently leads the FrontierSWE V2 leaderboard with a score of 0.623 across 5 evaluated AI models.
FAQ
Common questions about the FrontierSWE V2 benchmark and leaderboard.