The AI arena is free today

Open Superagent

USAMO25

Paper

Progress Over Time

Interactive timeline showing model performance evolution on USAMO25

State-of-the-art frontier
Open
Proprietary

USAMO25 Leaderboard

3 models
ContextCostLicense
1
2
3
Notice missing or incorrect data?
About this benchmark

What is USAMO25?

The 2025 United States of America Mathematical Olympiad (USAMO) benchmark consists of six challenging mathematical problems requiring rigorous proof-based reasoning. USAMO is the most prestigious high school mathematics competition in the United States, serving as the final round of the American Mathematics Competitions series. This benchmark evaluates models on mathematical problem-solving capabilities beyond simple numerical computation, focusing on formal mathematical reasoning and proof generation.

USAMO25 is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 1.0.

Compare leaders on the best AI for math and best AI for reasoning leaderboards.

Current leaders

Claude Mythos Preview from Anthropic currently leads the USAMO25 leaderboard with a score of 0.976 across 3 evaluated AI models.

1Claude Mythos PreviewAnthropic97.6%
2Grok-4 HeavyxAI61.9%
3Grok-4xAI37.5%

FAQ

Common questions about the USAMO25 benchmark and leaderboard.

What is the USAMO25 benchmark?

The 2025 United States of America Mathematical Olympiad (USAMO) benchmark consists of six challenging mathematical problems requiring rigorous proof-based reasoning. USAMO is the most prestigious high school mathematics competition in the United States, serving as the final round of the American Mathematics Competitions series. This benchmark evaluates models on mathematical problem-solving capabilities beyond simple numerical computation, focusing on formal mathematical reasoning and proof generation.

What is the USAMO25 leaderboard?

The USAMO25 leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Claude Mythos Preview by Anthropic leads with a score of 0.976. The average score across all models is 0.657.

What is the highest USAMO25 score?

The highest USAMO25 score is 0.976, achieved by Claude Mythos Preview from Anthropic.

How many models are evaluated on USAMO25?

3 models have been evaluated on the USAMO25 benchmark, with 0 verified results and 3 self-reported results.

Where can I find the USAMO25 paper?

The USAMO25 paper is available at https://arxiv.org/abs/2503.21934. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does USAMO25 cover?

USAMO25 is categorized under math and reasoning. The benchmark evaluates text models.

How recent are the USAMO25 leaderboard results?

The USAMO25 leaderboard was last updated in September 2026 and currently includes 3 evaluated models.