The AI arena is free today

Open Superagent

FrontierSWE V2

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on FrontierSWE V2

State-of-the-art frontier
Open
Proprietary

FrontierSWE V2 Leaderboard

5 models
ContextCostLicense
1—1.0M$4.00 / $20.00
2—1.0M$2.00 / $10.00
3———
4
Anthropic
Anthropic
—1.0M$0.10 / $0.50
5—500K$2.00 / $6.00
Notice missing or incorrect data?
About this benchmark

What is FrontierSWE V2?

FrontierSWE V2 evaluates agents on open-ended technical projects and reports mean@5. Kept separate from FrontierSWE V1 dominance-score results, which are not comparable.

FrontierSWE V2 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.6.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Claude Opus 5.5 from Anthropic currently leads the FrontierSWE V2 leaderboard with a score of 0.623 across 5 evaluated AI models.

1Claude Opus 5.5Anthropic62.3%
2Claude Sonnet 5.5Anthropic61.9%
3Gemini 4 ArgonGoogle55.0%

FAQ

Common questions about the FrontierSWE V2 benchmark and leaderboard.

What is the FrontierSWE V2 benchmark?

FrontierSWE V2 evaluates agents on open-ended technical projects and reports mean@5. Kept separate from FrontierSWE V1 dominance-score results, which are not comparable.

What is the FrontierSWE V2 leaderboard?

The FrontierSWE V2 leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, Claude Opus 5.5 by Anthropic leads with a score of 0.623. The average score across all models is 0.504.

What is the highest FrontierSWE V2 score?

The highest FrontierSWE V2 score is 0.623, achieved by Claude Opus 5.5 from Anthropic.

How many models are evaluated on FrontierSWE V2?

5 models have been evaluated on the FrontierSWE V2 benchmark, with 0 verified results and 5 self-reported results.

Where can I find the FrontierSWE V2 dataset?

The FrontierSWE V2 dataset is available at https://www.frontierswe.com/.

What categories does FrontierSWE V2 cover?

FrontierSWE V2 is categorized under agents and code. The benchmark evaluates text models.

What's the difference between FrontierSWE V2 and FrontierSWE?

FrontierSWE V2 is a variant of FrontierSWE. See the FrontierSWE leaderboard for the broader benchmark and per-model comparison.

Which model offers the best value on FrontierSWE V2?

Among models scoring within 10% of the leader, Claude Sonnet 5.5 from Anthropic is the cheapest, at $2.00 per million input tokens with a score of 0.619.

How recent are the FrontierSWE V2 leaderboard results?

The FrontierSWE V2 leaderboard was last updated in October 2026 and currently includes 5 evaluated models.