The AI arena is free today

Open Superagent

MTVQA

Paper

Progress Over Time

Interactive timeline showing model performance evolution on MTVQA

State-of-the-art frontier
Open
Proprietary

MTVQA Leaderboard

1 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
73B
Notice missing or incorrect data?
About this benchmark

What is MTVQA?

MTVQA (Multilingual Text-Centric Visual Question Answering) is the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. It addresses visual-textual misalignment problems in multilingual text-centric VQA.

MTVQA is a multimodal benchmark evaluating models on multimodal, text-to-image, and vision tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.3.

Compare leaders on the best AI for multimodal, best AI for text-to-image and best AI for vision leaderboards.

Current leaders

Qwen2-VL-72B-Instruct from Alibaba Cloud / Qwen Team currently leads the MTVQA leaderboard with a score of 0.309 across 1 evaluated AI models.

1Qwen2-VL-72B-InstructAlibaba Cloud / Qwen Team30.9%

FAQ

Common questions about the MTVQA benchmark and leaderboard.

What is the MTVQA benchmark?

MTVQA (Multilingual Text-Centric Visual Question Answering) is the first benchmark featuring high-quality human expert annotations across 9 diverse languages, consisting of 6,778 question-answer pairs across 2,116 images. It addresses visual-textual misalignment problems in multilingual text-centric VQA.

What is the MTVQA leaderboard?

The MTVQA leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Qwen2-VL-72B-Instruct by Alibaba Cloud / Qwen Team leads with a score of 0.309. The average score across all models is 0.309.

What is the highest MTVQA score?

The highest MTVQA score is 0.309, achieved by Qwen2-VL-72B-Instruct from Alibaba Cloud / Qwen Team.

How many models are evaluated on MTVQA?

1 models have been evaluated on the MTVQA benchmark, with 0 verified results and 1 self-reported results.

Where can I find the MTVQA paper?

The MTVQA paper is available at https://arxiv.org/abs/2405.11985. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MTVQA cover?

MTVQA is categorized under multimodal, text-to-image, and vision. The benchmark evaluates multimodal models with multilingual support.

What is the best open-source model on MTVQA?

Qwen2-VL-72B-Instruct by Alibaba Cloud / Qwen Team is the top-ranked open-source model on MTVQA, with a score of 0.309 (rank #1).

How recent are the MTVQA leaderboard results?

The MTVQA leaderboard was last updated in August 2026 and currently includes 1 evaluated models.