The AI arena is free today

Open Superagent

AIME 2026

Progress Over Time

Interactive timeline showing model performance evolution on AIME 2026

State-of-the-art frontier
Open
Proprietary

AIME 2026 Leaderboard

26 models
ContextCostLicense
1
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
2
Thinking Machines Lab
Thinking Machines Lab
975B524K$0.95 / $4.05
3
Sakana AI
Sakana AI
—256K$0.95 / $4.00
4
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
5
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
6—524K$0.30 / $1.20
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$0.50 / $3.00
6
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
930B131K$0.30 / $1.20
101.0T——
11
ByteDance
ByteDance
—256K$0.50 / $3.00
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
14
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
16
LG AI Research
LG AI Research
33B——
17———
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
1931B262K$0.09 / $0.34
20
ByteDance
ByteDance
———
2025B262K$0.07 / $0.34
22
ByteDance
ByteDance
—256K$0.10 / $0.40
2312B——
2425B——
258B131K$0.02 / $0.10
265B——
Notice missing or incorrect data?
About this benchmark

What is AIME 2026?

All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

AIME 2026 is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 26 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 1.0.

Compare leaders on the best AI for math and best AI for reasoning leaderboards.

Current leaders

GLM-5.2 from Zhipu AI currently leads the AIME 2026 leaderboard with a score of 0.992 across 26 evaluated AI models.

1GLM-5.2Zhipu AI99.2%
2InklingThinking Machines Lab97.1%
3Sakana NamazuSakana AI96.7%

FAQ

Common questions about the AIME 2026 benchmark and leaderboard.

What is the AIME 2026 benchmark?

All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.

What is the AIME 2026 leaderboard?

The AIME 2026 leaderboard ranks 26 AI models based on their performance on this benchmark. Currently, GLM-5.2 by Zhipu AI leads with a score of 0.992. The average score across all models is 0.878.

What is the highest AIME 2026 score?

The highest AIME 2026 score is 0.992, achieved by GLM-5.2 from Zhipu AI.

How many models are evaluated on AIME 2026?

26 models have been evaluated on the AIME 2026 benchmark, with 0 verified results and 26 self-reported results.

What categories does AIME 2026 cover?

AIME 2026 is categorized under math and reasoning. The benchmark evaluates text models.

What is the best open-source model on AIME 2026?

GLM-5.2 by Zhipu AI is the top-ranked open-source model on AIME 2026, with a score of 0.992 (rank #1).

Which model offers the best value on AIME 2026?

Among models scoring within 10% of the leader, Ling 3.0 Flash from InclusionAI is the cheapest, at $0.06 per million input tokens with a score of 0.932.

How recent are the AIME 2026 leaderboard results?

The AIME 2026 leaderboard was last updated in October 2026 and currently includes 26 evaluated models.