The AI arena is free today

Open Superagent

DocVQA

Paper

Progress Over Time

Interactive timeline showing model performance evolution on DocVQA

State-of-the-art frontier
Open
Proprietary

DocVQA Leaderboard

28 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
72B
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
7B
3
524B128K$0.07 / $0.20
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
34B
7400B1.0M$0.20 / $0.80
7109B328K$0.10 / $0.30
9
10
Amazon
Amazon
11
Mistral AI
Mistral AI
124B
11
DeepSeek
DeepSeek
27B
13
136B
15
OpenAI
OpenAI
128K$2.50 / $10.00
16
Amazon
Amazon
1716B
182B
19
Liquid AI
Liquid AI
3B
20
Mistral AI
Mistral AI
12B
2190B
223B
2311B
2412B131K$0.05 / $0.15
2527B131K$0.08 / $0.16
26
26
284B131K$0.05 / $0.10
Notice missing or incorrect data?
About this benchmark

What is DocVQA?

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

DocVQA is a multimodal benchmark evaluating models on image to text, multimodal, and vision tasks. LLM Stats tracks 28 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 1.0.

Compare leaders on the best AI for image to text, best AI for multimodal and best AI for vision leaderboards.

Current leaders

Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team currently leads the DocVQA leaderboard with a score of 0.964 across 28 evaluated AI models.

1Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team96.4%
2Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team95.7%
3Qwen2.5-Omni-7BAlibaba Cloud / Qwen Team95.2%

Source paper

Title
DocVQA: A Dataset for VQA on Document Images
Authors
Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar
Published
Abstract

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at docvqa.org

FAQ

Common questions about the DocVQA benchmark and leaderboard.

What is the DocVQA benchmark?

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

What is the DocVQA leaderboard?

The DocVQA leaderboard ranks 28 AI models based on their performance on this benchmark. Currently, Qwen2.5 VL 72B Instruct by Alibaba Cloud / Qwen Team leads with a score of 0.964. The average score across all models is 0.914.

What is the highest DocVQA score?

The highest DocVQA score is 0.964, achieved by Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team.

How many models are evaluated on DocVQA?

28 models have been evaluated on the DocVQA benchmark, with 0 verified results and 28 self-reported results.

Where can I find the DocVQA paper?

The DocVQA paper is available at https://arxiv.org/abs/2007.00398. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does DocVQA cover?

DocVQA is categorized under image to text, multimodal, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on DocVQA?

Qwen2.5 VL 7B Instruct by Alibaba Cloud / Qwen Team is the top-ranked open-source model on DocVQA, with a score of 0.957 (rank #2).

Which model offers the best value on DocVQA?

Among models scoring within 10% of the leader, Gemma 3 12B from Google is the cheapest, at $0.05 per million input tokens with a score of 0.871.

How recent are the DocVQA leaderboard results?

The DocVQA leaderboard was last updated in September 2026 and currently includes 28 evaluated models.