The AI arena is free today

Open Superagent

MME

Paper

Progress Over Time

Interactive timeline showing model performance evolution on MME

State-of-the-art frontier
Open
Proprietary

MME Leaderboard

4 models
ContextCostLicense
1
Liquid AI
Liquid AI
3B
2
DeepSeek
DeepSeek
27B
316B
43B
Notice missing or incorrect data?
About this benchmark

What is MME?

A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.

MME is a multimodal benchmark evaluating models on multimodal, reasoning, and vision tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.7.

Compare leaders on the best AI for multimodal, best AI for reasoning and best AI for vision leaderboards.

Current leaders

LFM2.5-VL-3B from Liquid AI currently leads the MME leaderboard with a score of 0.731 across 4 evaluated AI models.

1LFM2.5-VL-3BLiquid AI73.1%
2DeepSeek VL2DeepSeek22.5%
3DeepSeek VL2 SmallDeepSeek21.2%

Source paper

Title
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Authors
Chaoyou Fu, Peixian Chen, Yunhang Shen, Yulei Qin, and 10 others
Published
Abstract

Multimodal Large Language Model (MLLM) relies on the powerful LLM to perform multimodal tasks, showing amazing emergent abilities in recent studies, such as writing poems based on an image. However, it is difficult for these case studies to fully reflect the performance of MLLM, lacking a comprehensive evaluation. In this paper, we fill in this blank, presenting the first comprehensive MLLM Evaluation benchmark MME. It measures both perception and cognition abilities on a total of 14 subtasks. In order to avoid data leakage that may arise from direct use of public datasets for evaluation, the annotations of instruction-answer pairs are all manually designed. The concise instruction design allows us to fairly compare MLLMs, instead of struggling in prompt engineering. Besides, with such an instruction, we can also easily carry out quantitative statistics. A total of 30 advanced MLLMs are comprehensively evaluated on our MME, which not only suggests that existing MLLMs still have a large room for improvement, but also reveals the potential directions for the subsequent model optimization. The data are released at the project page https://github.com/BradyFU/Awesome-Multimodal-Large-Language-Models/tree/Evaluation.

FAQ

Common questions about the MME benchmark and leaderboard.

What is the MME benchmark?

A comprehensive evaluation benchmark for Multimodal Large Language Models measuring both perception and cognition abilities across 14 subtasks. Features manually designed instruction-answer pairs to avoid data leakage and provides systematic quantitative assessment of MLLM capabilities.

What is the MME leaderboard?

The MME leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, LFM2.5-VL-3B by Liquid AI leads with a score of 0.731. The average score across all models is 0.340.

What is the highest MME score?

The highest MME score is 0.731, achieved by LFM2.5-VL-3B from Liquid AI.

How many models are evaluated on MME?

4 models have been evaluated on the MME benchmark, with 0 verified results and 4 self-reported results.

Where can I find the MME paper?

The MME paper is available at https://arxiv.org/abs/2306.13394. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MME cover?

MME is categorized under multimodal, reasoning, and vision. The benchmark evaluates multimodal models.

How recent are the MME leaderboard results?

The MME leaderboard was last updated in August 2026 and currently includes 4 evaluated models.