The AI arena is free today

Open Superagent

MathVision

Paper

Progress Over Time

Interactive timeline showing model performance evolution on MathVision

State-of-the-art frontier
Open
Proprietary

MathVision Leaderboard

37 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B——
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B1.0M$0.15 / $0.47
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.40 / $3.00
5
ByteDance
ByteDance
———
6
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
7———
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
—1.0M$0.50 / $3.00
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B262K$0.29 / $2.40
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.26 / $2.60
1231B262K$0.09 / $0.34
13
Moonshot AI
Moonshot AI
1.0T——
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.14 / $1.00
1525B262K$0.07 / $0.34
16
ByteDance
ByteDance
—256K$0.25 / $2.00
1712B——
18
ByteDance
ByteDance
—256K$0.10 / $0.40
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B——
2010B——
2125B——
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B——
23
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B262K$0.20 / $0.88
24
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B——
25
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B——
26
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B——
27
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B262K$0.15 / $0.60
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B——
298B131K$0.02 / $0.10
30
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B——
315B——
32
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B——
33
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
34B——
34
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
72B——
35
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
73B——
36
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B——
37
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
7B——
Notice missing or incorrect data?
About this benchmark

What is MathVision?

MATH-Vision is a dataset designed to measure multimodal mathematical reasoning capabilities. It focuses on evaluating how well models can solve mathematical problems that require both visual understanding and mathematical reasoning, bridging the gap between visual and mathematical domains.

MathVision is a multimodal benchmark evaluating models on math, multimodal, and vision tasks. LLM Stats tracks 37 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 1.0.

Compare leaders on the best AI for math, best AI for multimodal and best AI for vision leaderboards.

Current leaders

Kimi K3 from Moonshot AI currently leads the MathVision leaderboard with a score of 0.978 across 37 evaluated AI models.

1Kimi K3Moonshot AI97.8%
2Qwen3.8-Flash-NextAlibaba Cloud / Qwen Team95.7%
2Qwen3.8 FlashAlibaba Cloud / Qwen Team95.7%
OSSQwen3.8-27B#4 open-weight94.6%

Source paper

Title
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
Authors
Ke Wang, Junting Pan, Weikang Shi, Zimu Lu, and 2 others
Published
Abstract

Recent advancements in Large Multimodal Models (LMMs) have shown promising results in mathematical reasoning within visual contexts, with models approaching human-level performance on existing benchmarks such as MathVista. However, we observe significant limitations in the diversity of questions and breadth of subjects covered by these benchmarks. To address this issue, we present the MATH-Vision (MATH-V) dataset, a meticulously curated collection of 3,040 high-quality mathematical problems with visual contexts sourced from real math competitions. Spanning 16 distinct mathematical disciplines and graded across 5 levels of difficulty, our dataset provides a comprehensive and diverse set of challenges for evaluating the mathematical reasoning abilities of LMMs. Through extensive experimentation, we unveil a notable performance gap between current LMMs and human performance on MATH-V, underscoring the imperative for further advancements in LMMs. Moreover, our detailed categorization allows for a thorough error analysis of LMMs, offering valuable insights to guide future research and development. The project is available at https://mathvision-cuhk.github.io

FAQ

Common questions about the MathVision benchmark and leaderboard.

What is the MathVision benchmark?

MATH-Vision is a dataset designed to measure multimodal mathematical reasoning capabilities. It focuses on evaluating how well models can solve mathematical problems that require both visual understanding and mathematical reasoning, bridging the gap between visual and mathematical domains.

What is the MathVision leaderboard?

The MathVision leaderboard ranks 37 AI models based on their performance on this benchmark. Currently, Kimi K3 by Moonshot AI leads with a score of 0.978. The average score across all models is 0.712.

What is the highest MathVision score?

The highest MathVision score is 0.978, achieved by Kimi K3 from Moonshot AI.

How many models are evaluated on MathVision?

37 models have been evaluated on the MathVision benchmark, with 0 verified results and 37 self-reported results.

Where can I find the MathVision paper?

The MathVision paper is available at https://arxiv.org/abs/2402.14804. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MathVision cover?

MathVision is categorized under math, multimodal, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on MathVision?

Qwen3.8-27B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on MathVision, with a score of 0.946 (rank #4).

Which model offers the best value on MathVision?

Among models scoring within 10% of the leader, Qwen3.8 Flash from Alibaba Cloud / Qwen Team is the cheapest, at $0.15 per million input tokens with a score of 0.957.

How recent are the MathVision leaderboard results?

The MathVision leaderboard was last updated in October 2026 and currently includes 37 evaluated models.