The AI arena is free today

Open Superagent

Agents' Last Exam

Progress Over Time

Interactive timeline showing model performance evolution on Agents' Last Exam

State-of-the-art frontier
Open
Proprietary

Agents' Last Exam Leaderboard

21 models
ContextCostLicense
1—1.1M$10.00 / $50.00
2
OpenAI
OpenAI
—1.1M$2.00 / $10.00
3—1.1M$5.00 / $30.00
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B——
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
7
OpenAI
OpenAI
—1.1M$0.10 / $0.50
8—1.1M$2.00 / $12.00
9—1.1M$0.20 / $1.20
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.40 / $3.00
11
ByteDance
ByteDance
———
12763B1.0M$0.22 / $0.66
13
Xiaomi
Xiaomi
1.0T1.0M$0.43 / $0.87
14
Zhipu AI
Zhipu AI
753B1.0M$1.20 / $4.00
15309B1.0M$0.14 / $0.28
16—1.0M$0.44 / $1.32
17—1.0M$0.75 / $3.75
17320B1.0M$0.15 / $0.50
191.6T1.0M$1.30 / $2.60
20304B1.0M$0.06 / $0.18
21770B——
Notice missing or incorrect data?
About this benchmark

What is Agents' Last Exam?

Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access.

Agents' Last Exam is a text benchmark evaluating models on reasoning, agents, and tool calling tasks. LLM Stats tracks 21 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.6.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for tool calling leaderboards.

Current leaders

GPT-6 Astra from OpenAI currently leads the Agents' Last Exam leaderboard with a score of 0.593 across 21 evaluated AI models.

1GPT-6 AstraOpenAI59.3%
2GPT-6 SolOpenAI56.4%
3GPT-5.6 SolOpenAI52.7%
OSSQwen3.8-27B#10 open-weight42.9%

FAQ

Common questions about the Agents' Last Exam benchmark and leaderboard.

What is the Agents' Last Exam benchmark?

Agents' Last Exam is a challenging benchmark for AI agents on hard, long-horizon tasks that test sustained reasoning, planning, and tool use, reported with and without tool access.

What is the Agents' Last Exam leaderboard?

The Agents' Last Exam leaderboard ranks 21 AI models based on their performance on this benchmark. Currently, GPT-6 Astra by OpenAI leads with a score of 0.593. The average score across all models is 0.396.

What is the highest Agents' Last Exam score?

The highest Agents' Last Exam score is 0.593, achieved by GPT-6 Astra from OpenAI.

How many models are evaluated on Agents' Last Exam?

21 models have been evaluated on the Agents' Last Exam benchmark, with 0 verified results and 21 self-reported results.

What categories does Agents' Last Exam cover?

Agents' Last Exam is categorized under reasoning, agents, and tool calling. The benchmark evaluates text models.

What is the best open-source model on Agents' Last Exam?

Qwen3.8-27B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on Agents' Last Exam, with a score of 0.429 (rank #10).

Which model offers the best value on Agents' Last Exam?

Among models scoring within 10% of the leader, GPT-6 Sol from OpenAI is the cheapest, at $2.00 per million input tokens with a score of 0.564.

How recent are the Agents' Last Exam leaderboard results?

The Agents' Last Exam leaderboard was last updated in September 2026 and currently includes 21 evaluated models.