The AI arena is free today

Open Superagent

IMO-AnswerBench

Progress Over Time

Interactive timeline showing model performance evolution on IMO-AnswerBench

State-of-the-art frontier
Open
Proprietary

IMO-AnswerBench Leaderboard

22 models
ContextCostLicense
1550B262K$0.50 / $2.20
2
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
3
Tencent
Tencent
295B262K$0.14 / $0.58
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$1.25 / $3.75
51.6T1.0M$1.30 / $2.60
6284B1.0M$0.09 / $0.18
7
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
9196B66K$0.10 / $0.40
10284B1.0M$0.09 / $0.18
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0T256K$1.20 / $6.00
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$0.50 / $3.00
12
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
14
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
15
Zhipu AI
Zhipu AI
358B203K$0.40 / $1.75
16
Moonshot AI
Moonshot AI
1.0T——
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B262K$0.45 / $3.00
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
20560B——
201.0T——
22685B164K$0.26 / $0.38
Notice missing or incorrect data?
About this benchmark

What is IMO-AnswerBench?

IMO-AnswerBench is a benchmark for evaluating mathematical reasoning capabilities on International Mathematical Olympiad (IMO) problems, focusing on answer generation and verification.

IMO-AnswerBench is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 22 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for math and best AI for reasoning leaderboards.

Current leaders

Nemotron 3 Ultra (550B A55B) from NVIDIA currently leads the IMO-AnswerBench leaderboard with a score of 0.923 across 22 evaluated AI models.

2GLM-5.2Zhipu AI91.0%
3Hy3Tencent90.0%

FAQ

Common questions about the IMO-AnswerBench benchmark and leaderboard.

What is the IMO-AnswerBench benchmark?

IMO-AnswerBench is a benchmark for evaluating mathematical reasoning capabilities on International Mathematical Olympiad (IMO) problems, focusing on answer generation and verification.

What is the IMO-AnswerBench leaderboard?

The IMO-AnswerBench leaderboard ranks 22 AI models based on their performance on this benchmark. Currently, Nemotron 3 Ultra (550B A55B) by NVIDIA leads with a score of 0.923. The average score across all models is 0.845.

What is the highest IMO-AnswerBench score?

The highest IMO-AnswerBench score is 0.923, achieved by Nemotron 3 Ultra (550B A55B) from NVIDIA.

How many models are evaluated on IMO-AnswerBench?

22 models have been evaluated on the IMO-AnswerBench benchmark, with 0 verified results and 22 self-reported results.

What categories does IMO-AnswerBench cover?

IMO-AnswerBench is categorized under math and reasoning. The benchmark evaluates text models.

What is the best open-source model on IMO-AnswerBench?

Nemotron 3 Ultra (550B A55B) by NVIDIA is the top-ranked open-source model on IMO-AnswerBench, with a score of 0.923 (rank #1).

Which model offers the best value on IMO-AnswerBench?

Among models scoring within 10% of the leader, Ling 3.0 Flash from InclusionAI is the cheapest, at $0.06 per million input tokens with a score of 0.837.

How recent are the IMO-AnswerBench leaderboard results?

The IMO-AnswerBench leaderboard was last updated in October 2026 and currently includes 22 evaluated models.