AIME 2026
Progress Over Time
Interactive timeline showing model performance evolution on AIME 2026
AIME 2026 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Zhipu AI | 753B | 1.0M | $0.75 / $2.40 | ||
| 2 | Thinking Machines Lab | 975B | 524K | $0.95 / $4.05 | ||
| 3 | Sakana AI | — | 256K | $0.95 / $4.00 | ||
| 4 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 5 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 | ||
| 6 | Upstage | — | 524K | $0.30 / $1.20 | ||
| 6 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 6 | Zhipu AI | 754B | 203K | $1.05 / $3.50 | ||
| 9 | Meta | 30B | 131K | $0.30 / $1.20 | ||
| 10 | Microsoft | 1.0T | — | — | ||
| 11 | ByteDance | — | 256K | $0.50 / $3.00 | ||
| 12 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.32 / $3.20 | ||
| 13 | Alibaba Cloud / Qwen Team | 1.0T | 256K | $1.20 / $6.00 | ||
| 14 | InclusionAI | 124B | 131K | $0.06 / $0.18 | ||
| 15 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.10 / $0.95 | ||
| 16 | LG AI Research | 33B | — | — | ||
| 17 | Microsoft | — | — | — | ||
| 18 | Alibaba Cloud / Qwen Team | 397B | 262K | $0.45 / $3.00 | ||
| 19 | Google | 31B | 262K | $0.09 / $0.34 | ||
| 20 | ByteDance | — | — | — | ||
| 20 | Google | 25B | 262K | $0.07 / $0.34 | ||
| 22 | ByteDance | — | 256K | $0.10 / $0.40 | ||
| 23 | Google | 12B | — | — | ||
| 24 | Google | 25B | — | — | ||
| 25 | Google | 8B | 131K | $0.02 / $0.10 | ||
| 26 | Google | 5B | — | — |
What is AIME 2026?
All 30 problems from the 2026 American Invitational Mathematics Examination (AIME I and AIME II), testing olympiad-level mathematical reasoning with integer answers from 000-999. Used as an AI benchmark to evaluate large language models' ability to solve complex mathematical problems requiring multi-step logical deductions and structured symbolic reasoning.
AIME 2026 is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 26 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 1.0.
Compare leaders on the best AI for math and best AI for reasoning leaderboards.
Current leaders
GLM-5.2 from Zhipu AI currently leads the AIME 2026 leaderboard with a score of 0.992 across 26 evaluated AI models.
FAQ
Common questions about the AIME 2026 benchmark and leaderboard.