FrontierCode 1.1

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on FrontierCode 1.1

State-of-the-art frontier
Open
Proprietary

FrontierCode 1.1 Leaderboard

15 models
ContextCostLicense
11.0M$10.00 / $50.00
2
Anthropic
Anthropic
1.0M$5.00 / $25.00
31.1M$5.00 / $30.00
41.0M$5.00 / $25.00
5
OpenAI
OpenAI
1.1M$5.00 / $30.00
61.0M$3.00 / $15.00
7500K$2.00 / $6.00
81.1M$2.50 / $15.00
91.1M$1.00 / $6.00
101.0M$5.00 / $25.00
11
Moonshot AI
Moonshot AI
1.0T262K$0.74 / $3.50
12
Zhipu AI
Zhipu AI
753B1.0M$0.95 / $3.00
131.6T1.0M$1.60 / $3.20
14
MiniMax
MiniMax
1.0M$0.30 / $1.20
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.32 / $1.28
Notice missing or incorrect data?
About this benchmark

What is FrontierCode 1.1?

FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.

FrontierCode 1.1 is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 15 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.

Current leaders

Claude Fable 5 from Anthropic currently leads the FrontierCode 1.1 leaderboard with a score of 0.535 across 15 evaluated AI models.

1Claude Fable 5Anthropic53.5%
2Claude Opus 5Anthropic53.4%
3GPT-5.6 SolOpenAI47.5%
OSSKimi K2.7 Code#11 open-weight30.1%

FAQ

Common questions about the FrontierCode 1.1 benchmark and leaderboard.

What is the FrontierCode 1.1 benchmark?

FrontierCode 1.1 evaluates whether coding-agent changes are mergeable, using unit tests, maintainer-defined rubrics, and verifiers. Runs flagged for unfair internet use receive a zero score.

What is the FrontierCode 1.1 leaderboard?

The FrontierCode 1.1 leaderboard ranks 15 AI models based on their performance on this benchmark. Currently, Claude Fable 5 by Anthropic leads with a score of 0.535. The average score across all models is 0.364.

What is the highest FrontierCode 1.1 score?

The highest FrontierCode 1.1 score is 0.535, achieved by Claude Fable 5 from Anthropic.

How many models are evaluated on FrontierCode 1.1?

15 models have been evaluated on the FrontierCode 1.1 benchmark, with 0 verified results and 0 self-reported results.

Where can I find the FrontierCode 1.1 dataset?

The FrontierCode 1.1 dataset is available at https://cognition.com/frontiercode.

What categories does FrontierCode 1.1 cover?

FrontierCode 1.1 is categorized under reasoning, agents, and code. The benchmark evaluates text models.

What's the difference between FrontierCode 1.1 and FrontierCode?

FrontierCode 1.1 is a variant of FrontierCode. See the FrontierCode leaderboard for the broader benchmark and per-model comparison.

What is the best open-source model on FrontierCode 1.1?

Kimi K2.7 Code by Moonshot AI is the top-ranked open-source model on FrontierCode 1.1, with a score of 0.301 (rank #11).

Which model offers the best value on FrontierCode 1.1?

Among models scoring within 10% of the leader, Claude Opus 5 from Anthropic is the cheapest, at $5.00 per million input tokens with a score of 0.534.

How recent are the FrontierCode 1.1 leaderboard results?

The FrontierCode 1.1 leaderboard was last updated in July 2026 and currently includes 15 evaluated models.