MMMU (validation)
Progress Over Time
Interactive timeline showing model performance evolution on MMMU (validation)
MMMU (validation) Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | — | — | ||
| 2 | Anthropic | — | — | — | ||
| 3 | Anthropic | — | — | — | ||
| 4 | Anthropic | — | 200K | $1.00 / $5.00 |
What is MMMU (validation)?
Validation set of the Massive Multi-discipline Multimodal Understanding and Reasoning benchmark. Features college-level multimodal questions across 6 core disciplines (Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, Tech & Engineering) spanning 30 subjects and 183 subfields with diverse image types including charts, diagrams, maps, and tables.
MMMU (validation) is a multimodal benchmark evaluating models on multimodal, reasoning, general, healthcare, and vision tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for multimodal, best AI for reasoning, best AI for general, best AI for healthcare and best AI for vision leaderboards.
Current leaders
Claude Opus 4.5 from Anthropic currently leads the MMMU (validation) leaderboard with a score of 0.807 across 4 evaluated AI models.
Source paper
- Title
- MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
- Authors
- Xiang Yue, Yuansheng Ni, Kai Zhang, Tianyu Zheng, and 18 others
- Published
- arXiv
- 2311.16502
Abstract
We introduce MMMU: a new benchmark designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. MMMU includes 11.5K meticulously collected multimodal questions from college exams, quizzes, and textbooks, covering six core disciplines: Art & Design, Business, Science, Health & Medicine, Humanities & Social Science, and Tech & Engineering. These questions span 30 subjects and 183 subfields, comprising 30 highly heterogeneous image types, such as charts, diagrams, maps, tables, music sheets, and chemical structures. Unlike existing benchmarks, MMMU focuses on advanced perception and reasoning with domain-specific knowledge, challenging models to perform tasks akin to those faced by experts. The evaluation of 14 open-source LMMs as well as the proprietary GPT-4V(ision) and Gemini highlights the substantial challenges posed by MMMU. Even the advanced GPT-4V and Gemini Ultra only achieve accuracies of 56% and 59% respectively, indicating significant room for improvement. We believe MMMU will stimulate the community to build next-generation multimodal foundation models towards expert artificial general intelligence.
FAQ
Common questions about the MMMU (validation) benchmark and leaderboard.