The AI arena is free today

Open Superagent

MMMU-Pro

Progress Over Time

Interactive timeline showing model performance evolution on MMMU-Pro

State-of-the-art frontier
Open
Proprietary

MMMU-Pro Leaderboard

72 models
ContextCostLicense
1—1.0M$1.50 / $9.00
2
OpenAI
OpenAI
—1.1M$5.00 / $30.00
3—1.1M$5.00 / $30.00
4
ByteDance
ByteDance
———
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
6———
7
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
8—1.0M$0.50 / $3.00
8
OpenAI
OpenAI
—1.0M$2.50 / $15.00
10———
11—1.1M$2.00 / $12.00
12—1.0M$2.00 / $12.00
13———
14
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
15
OpenAI
OpenAI
—400K$1.75 / $14.00
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$0.50 / $3.00
18
Moonshot AI
Moonshot AI
1.0T——
19
OpenAI
OpenAI
———
19—1.1M$0.20 / $1.20
21
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
22
Xiaomi
Xiaomi
311B1.0M$0.17 / $0.34
23—1.0M$5.00 / $25.00
2431B262K$0.09 / $0.34
24
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
26—1.0M$0.25 / $1.50
27—400K$0.75 / $4.50
28
OpenAI
OpenAI
———
29—400K$5.00 / $30.00
30
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.32 / $3.20
31—1.0M$3.00 / $15.00
32
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
33
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
34
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
3530B131K$0.30 / $1.20
35
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
3725B262K$0.07 / $0.34
38
Thinking Machines Lab
Thinking Machines Lab
975B524K$0.95 / $4.05
39
ByteDance
ByteDance
—256K$0.25 / $2.00
40
ByteDance
ByteDance
—256K$0.10 / $0.40
41
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B——
4212B——
43
LG AI Research
LG AI Research
33B——
44
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B262K$0.20 / $0.88
44
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B——
46—400K$0.20 / $1.25
47
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B——
48———
49
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B——
49218B——
1–50 of 72
1/2
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is MMMU-Pro?

A more robust multi-discipline multimodal understanding benchmark that enhances MMMU through a three-step process: filtering text-only answerable questions, augmenting candidate options, and introducing vision-only input settings. Achieves significantly lower model performance (16.8-26.9%) compared to original MMMU, providing more rigorous evaluation that closely mimics real-world scenarios.

MMMU-Pro is a multimodal benchmark evaluating models on multimodal, reasoning, general, and vision tasks. LLM Stats tracks 72 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.

Compare leaders on the best AI for multimodal, best AI for reasoning, best AI for general and best AI for vision leaderboards.

Current leaders

Gemini 3.5 Flash from Google currently leads the MMMU-Pro leaderboard with a score of 0.836 across 72 evaluated AI models.

1Gemini 3.5 FlashGoogle83.6%
2GPT-5.5OpenAI83.2%
3GPT-5.6 SolOpenAI83.0%
OSSKimi K2.5#18 open-weight78.5%

Source paper

Title
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Authors
Xiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang, and 9 others
Published
Abstract

This paper introduces MMMU-Pro, a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark. MMMU-Pro rigorously assesses multimodal models' true understanding and reasoning capabilities through a three-step process based on MMMU: (1) filtering out questions answerable by text-only models, (2) augmenting candidate options, and (3) introducing a vision-only input setting where questions are embedded within images. This setting challenges AI to truly "see" and "read" simultaneously, testing a fundamental human cognitive skill of seamlessly integrating visual and textual information. Results show that model performance is substantially lower on MMMU-Pro than on MMMU, ranging from 16.8% to 26.9% across models. We explore the impact of OCR prompts and Chain of Thought (CoT) reasoning, finding that OCR prompts have minimal effect while CoT generally improves performance. MMMU-Pro provides a more rigorous evaluation tool, closely mimicking real-world scenarios and offering valuable directions for future research in multimodal AI.

FAQ

Common questions about the MMMU-Pro benchmark and leaderboard.

What is the MMMU-Pro benchmark?

A more robust multi-discipline multimodal understanding benchmark that enhances MMMU through a three-step process: filtering text-only answerable questions, augmenting candidate options, and introducing vision-only input settings. Achieves significantly lower model performance (16.8-26.9%) compared to original MMMU, providing more rigorous evaluation that closely mimics real-world scenarios.

What is the MMMU-Pro leaderboard?

The MMMU-Pro leaderboard ranks 72 AI models based on their performance on this benchmark. Currently, Gemini 3.5 Flash by Google leads with a score of 0.836. The average score across all models is 0.680.

What is the highest MMMU-Pro score?

The highest MMMU-Pro score is 0.836, achieved by Gemini 3.5 Flash from Google.

How many models are evaluated on MMMU-Pro?

72 models have been evaluated on the MMMU-Pro benchmark, with 0 verified results and 72 self-reported results.

Where can I find the MMMU-Pro paper?

The MMMU-Pro paper is available at https://arxiv.org/abs/2409.02813. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MMMU-Pro cover?

MMMU-Pro is categorized under multimodal, reasoning, general, and vision. The benchmark evaluates multimodal models.

Are there variants of MMMU-Pro?

Yes. MMMU-Pro has 1 related variant: MMMU-Pro (with tools).

What is the best open-source model on MMMU-Pro?

Kimi K2.5 by Moonshot AI is the top-ranked open-source model on MMMU-Pro, with a score of 0.785 (rank #18).

Which model offers the best value on MMMU-Pro?

Among models scoring within 10% of the leader, Gemma 4 31B from Google is the cheapest, at $0.09 per million input tokens with a score of 0.769.

How is MMMU-Pro scored?

MMMU-Pro is scored using accuracy, reported on a 0–1 scale. Lower is better only when explicitly noted; on this leaderboard, higher scores indicate better performance.

How recent are the MMMU-Pro leaderboard results?

The MMMU-Pro leaderboard was last updated in October 2026 and currently includes 72 evaluated models.