DocVQA

Paper

Progress Over Time

Interactive timeline showing model performance evolution on DocVQA

State-of-the-art frontier
Open
Proprietary

DocVQA Leaderboard

26 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
72B
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
3
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
7B
524B
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
34B
7109B
7400B
9
10
Amazon
Amazon
11
DeepSeek
DeepSeek
27B
11
Mistral AI
Mistral AI
124B
136B
13
15
OpenAI
OpenAI
128K$2.50 / $10.00
16
Amazon
Amazon
1716B
18
Mistral AI
Mistral AI
12B
1990B
203B
2111B
2212B
2327B
24
24
264B
Notice missing or incorrect data?
About this benchmark

What is DocVQA?

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

DocVQA is a multimodal benchmark evaluating models on multimodal, image to text, and vision tasks. LLM Stats tracks 26 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 1.0.

Compare leaders on the best AI for multimodal, best AI for image to text and best AI for vision leaderboards.

Current leaders

Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team currently leads the DocVQA leaderboard with a score of 0.964 across 26 evaluated AI models.

1Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team96.4%
2Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team95.7%
3Claude 3.5 SonnetAnthropic95.2%

Source paper

Title
DocVQA: A Dataset for VQA on Document Images
Authors
Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar
Published
Abstract

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed analysis of the dataset in comparison with similar datasets for VQA and reading comprehension is presented. We report several baseline results by adopting existing VQA and reading comprehension models. Although the existing models perform reasonably well on certain types of questions, there is large performance gap compared to human performance (94.36% accuracy). The models need to improve specifically on questions where understanding structure of the document is crucial. The dataset, code and leaderboard are available at docvqa.org

FAQ

Common questions about the DocVQA benchmark and leaderboard.

What is the DocVQA benchmark?

A dataset for Visual Question Answering on document images containing 50,000 questions defined on 12,000+ document images. The benchmark tests AI's ability to understand document structure and content, requiring models to comprehend document layout and perform information retrieval to answer questions about document images.

What is the DocVQA leaderboard?

The DocVQA leaderboard ranks 26 AI models based on their performance on this benchmark. Currently, Qwen2.5 VL 72B Instruct by Alibaba Cloud / Qwen Team leads with a score of 0.964. The average score across all models is 0.914.

What is the highest DocVQA score?

The highest DocVQA score is 0.964, achieved by Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team.

How many models are evaluated on DocVQA?

26 models have been evaluated on the DocVQA benchmark, with 0 verified results and 26 self-reported results.

Where can I find the DocVQA paper?

The DocVQA paper is available at https://arxiv.org/abs/2007.00398. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does DocVQA cover?

DocVQA is categorized under multimodal, image to text, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on DocVQA?

Qwen2.5 VL 72B Instruct by Alibaba Cloud / Qwen Team is the top-ranked open-source model on DocVQA, with a score of 0.964 (rank #1).

Which model offers the best value on DocVQA?

Among models scoring within 10% of the leader, GPT-4o from OpenAI is the cheapest, at $2.50 per million input tokens with a score of 0.928.

How recent are the DocVQA leaderboard results?

The DocVQA leaderboard was last updated in August 2026 and currently includes 26 evaluated models.