MLVU-M
Progress Over Time
Interactive timeline showing model performance evolution on MLVU-M
State-of-the-art frontier
Open
Proprietary
MLVU-M Leaderboard
8 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 3 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $1.00 | ||
| 6 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $0.60 | ||
| 7 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 72B | — | — |
Notice missing or incorrect data?
What is MLVU-M?
MLVU-M benchmark
MLVU-M is a text benchmark evaluating models on general tasks. LLM Stats tracks 8 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for general leaderboards.
Current leaders
Qwen3 VL 32B Instruct from Alibaba Cloud / Qwen Team currently leads the MLVU-M leaderboard with a score of 0.821 across 8 evaluated AI models.
FAQ
Common questions about the MLVU-M benchmark and leaderboard.