FrontierCode
Progress Over Time
Interactive timeline showing model performance evolution on FrontierCode
FrontierCode Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 2 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 3 | Anthropic | — | 1.0M | $3.00 / $15.00 |
Sub-benchmarks
What is FrontierCode?
FrontierCode is Cognition's coding evaluation that tests whether models can pass difficult coding tasks while meeting the standards of high-quality production codebases. The Diamond subset contains the hardest problems.
FrontierCode is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.
Current leaders
Claude Opus 5 from Anthropic currently leads the FrontierCode leaderboard with a score of 0.534 across 3 evaluated AI models.
FAQ
Common questions about the FrontierCode benchmark and leaderboard.