OlympiadBench
Progress Over Time
Interactive timeline showing model performance evolution on OlympiadBench
OlympiadBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 73B | — | — |
What is OlympiadBench?
A challenging benchmark for promoting AGI with Olympiad-level bilingual multimodal scientific problems. Comprises 8,476 math and physics problems from international and Chinese Olympiads and the Chinese college entrance exam, featuring expert-level annotations for step-by-step reasoning. Includes both text-only and multimodal problems in English and Chinese.
OlympiadBench is a multimodal benchmark evaluating models on math, multimodal, physics, reasoning, and vision tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.2.
Compare leaders on the best AI for math, best AI for multimodal, best AI for physics, best AI for reasoning and best AI for vision leaderboards.
Current leaders
QvQ-72B-Preview from Alibaba Cloud / Qwen Team currently leads the OlympiadBench leaderboard with a score of 0.204 across 1 evaluated AI models.
FAQ
Common questions about the OlympiadBench benchmark and leaderboard.