FrontierCode

Progress Over Time

Interactive timeline showing model performance evolution on FrontierCode

State-of-the-art frontier
Open
Proprietary

FrontierCode Leaderboard

3 models
ContextCostLicense
1
Anthropic
Anthropic
1.0M$5.00 / $25.00
21.0M$10.00 / $50.00
31.0M$3.00 / $15.00
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is FrontierCode?

FrontierCode is Cognition's coding evaluation that tests whether models can pass difficult coding tasks while meeting the standards of high-quality production codebases. The Diamond subset contains the hardest problems.

FrontierCode is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.

Current leaders

Claude Opus 5 from Anthropic currently leads the FrontierCode leaderboard with a score of 0.534 across 3 evaluated AI models.

1Claude Opus 5Anthropic53.4%
2Claude Fable 5Anthropic46.3%
3Claude Sonnet 5Anthropic38.8%

FAQ

Common questions about the FrontierCode benchmark and leaderboard.

What is the FrontierCode benchmark?

FrontierCode is Cognition's coding evaluation that tests whether models can pass difficult coding tasks while meeting the standards of high-quality production codebases. The Diamond subset contains the hardest problems.

What is the FrontierCode leaderboard?

The FrontierCode leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Claude Opus 5 by Anthropic leads with a score of 0.534. The average score across all models is 0.462.

What is the highest FrontierCode score?

The highest FrontierCode score is 0.534, achieved by Claude Opus 5 from Anthropic.

How many models are evaluated on FrontierCode?

3 models have been evaluated on the FrontierCode benchmark, with 0 verified results and 3 self-reported results.

What categories does FrontierCode cover?

FrontierCode is categorized under reasoning, agents, and code. The benchmark evaluates text models.

Are there variants of FrontierCode?

Yes. FrontierCode has 1 related variant: FrontierCode 1.1.

Which model offers the best value on FrontierCode?

Among models scoring within 10% of the leader, Claude Opus 5 from Anthropic is the cheapest, at $5.00 per million input tokens with a score of 0.534.

How recent are the FrontierCode leaderboard results?

The FrontierCode leaderboard was last updated in July 2026 and currently includes 3 evaluated models.