MMLU-Redux
Progress Over Time
Interactive timeline showing model performance evolution on MMLU-Redux
MMLU-Redux Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 2 | Alibaba Cloud / Qwen Team | 397B | 262K | $0.45 / $3.00 | ||
| 3 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 3 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 5 | Moonshot AI | 1.0T | — | — | ||
| 6 | Alibaba Cloud / Qwen Team | 122B | 262K | $0.29 / $2.40 | ||
| 7 | Alibaba Cloud / Qwen Team | 235B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 236B | — | — | ||
| 9 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.32 / $3.20 | ||
| 10 | DeepSeek | 671B | 164K | $0.50 / $2.15 | ||
| 11 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.10 / $0.95 | ||
| 11 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.14 / $1.00 | ||
| 13 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.26 / $2.60 | ||
| 14 | Alibaba Cloud / Qwen Team | 235B | 262K | $0.09 / $0.55 | ||
| 15 | Alibaba Cloud / Qwen Team | 1.0T | 256K | $1.20 / $6.00 | ||
| 15 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 | ||
| 17 | Moonshot AI | 1.0T | — | — | ||
| 17 | Moonshot AI | 1.0T | — | — | ||
| 19 | Alibaba Cloud / Qwen Team | 80B | — | — | ||
| 20 | Alibaba Cloud / Qwen Team | 236B | 262K | $0.20 / $0.88 | ||
| 21 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 22 | DeepSeek | 671B | 164K | $0.25 / $0.95 | ||
| 23 | Alibaba Cloud / Qwen Team | 9B | 262K | $0.10 / $0.15 | ||
| 24 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 24 | Alibaba Cloud / Qwen Team | 80B | 262K | $0.09 / $1.10 | ||
| 26 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 27 | Meituan | 560B | — | — | ||
| 28 | DeepSeek | 671B | 164K | $0.32 / $0.89 | ||
| 29 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 29 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 31 | Alibaba Cloud / Qwen Team | 14B | 41K | $0.12 / $0.24 | ||
| 32 | Alibaba Cloud / Qwen Team | 31B | 262K | $0.15 / $0.60 | ||
| 33 | Alibaba Cloud / Qwen Team | 235B | — | — | ||
| 34 | Alibaba Cloud / Qwen Team | 73B | 33K | $0.36 / $0.40 | ||
| 35 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 36 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 37 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 38 | Mistral AI | 675B | — | — | ||
| 38 | Mistral AI | 14B | — | — | ||
| 40 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 41 | Alibaba Cloud / Qwen Team | 15B | — | — | ||
| 42 | Alibaba Cloud / Qwen Team | 2B | — | — | ||
| 43 | Mistral AI | 8B | — | — | ||
| 44 | Alibaba Cloud / Qwen Team | 32B | — | — | ||
| 45 | Alibaba Cloud / Qwen Team | 8B | — | — | ||
| 46 | Mistral AI | 3B | — | — | ||
| 47 | Alibaba Cloud / Qwen Team | 7B | — | — | ||
| 48 | Alibaba Cloud / Qwen Team | 7B | — | — | ||
| 49 | Alibaba Cloud / Qwen Team | 800M | — | — | ||
| 50 | Baidu | 21B | — | — |
What is MMLU-Redux?
An improved version of the MMLU benchmark featuring manually re-annotated questions to identify and correct errors in the original dataset. Provides more reliable evaluation metrics for language models by addressing dataset quality issues found in the original MMLU.
MMLU-Redux is a text benchmark evaluating models on language, math, reasoning, and general tasks. LLM Stats tracks 50 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.
Compare leaders on the best AI for language, best AI for math, best AI for reasoning and best AI for general leaderboards.
Current leaders
Qwen3.7 Max from Alibaba Cloud / Qwen Team currently leads the MMLU-Redux leaderboard with a score of 0.950 across 50 evaluated AI models.
Source paper
- Title
- Are We Done with MMLU?
- Authors
- Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, and 12 others
- Published
- arXiv
- 2406.04127
Abstract
Maybe not. We identify and analyse errors in the popular Massive Multitask Language Understanding (MMLU) benchmark. Even though MMLU is widely adopted, our analysis demonstrates numerous ground truth errors that obscure the true capabilities of LLMs. For example, we find that 57% of the analysed questions in the Virology subset contain errors. To address this issue, we introduce a comprehensive framework for identifying dataset errors using a novel error annotation protocol. Then, we create MMLU-Redux, which is a subset of 5,700 manually re-annotated questions across all 57 MMLU subjects. We estimate that 6.49% of MMLU questions contain errors. Using MMLU-Redux, we demonstrate significant discrepancies with the model performance metrics that were originally reported. Our results strongly advocate for revising MMLU's error-ridden questions to enhance its future utility and reliability as a benchmark. https://huggingface.co/datasets/edinburgh-dawg/mmlu-redux-2.0.
FAQ
Common questions about the MMLU-Redux benchmark and leaderboard.