FrontierCode 1.1
Progress Over Time
Interactive timeline showing model performance evolution on FrontierCode 1.1
FrontierCode 1.1 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $10.00 / $50.00 | ||
| 2 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 3 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 4 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 5 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 6 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 7 | xAI | — | 500K | $2.00 / $6.00 | ||
| 8 | OpenAI | — | 1.1M | $2.50 / $15.00 | ||
| 9 | OpenAI | — | 1.1M | $1.00 / $6.00 | ||
| 10 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 11 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 | ||
| 12 | Zhipu AI | 753B | 1.0M | $0.95 / $3.00 | ||
| 13 | DeepSeek | 1.6T | 1.0M | $1.60 / $3.20 | ||
| 14 | MiniMax | — | 1.0M | $0.30 / $1.20 | ||
| 15 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.32 / $1.28 |
What is FrontierCode 1.1?
FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.
FrontierCode 1.1 is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 15 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.
Current leaders
Claude Fable 5 from Anthropic currently leads the FrontierCode 1.1 leaderboard with a score of 0.535 across 15 evaluated AI models.
FAQ
Common questions about the FrontierCode 1.1 benchmark and leaderboard.