The AI arena is free today

Open Superagent

VQA-Rad

Paper

Progress Over Time

Interactive timeline showing model performance evolution on VQA-Rad

State-of-the-art frontier
Open
Proprietary

VQA-Rad Leaderboard

1 models
ContextCostLicense
14B
Notice missing or incorrect data?
About this benchmark

What is VQA-Rad?

VQA-RAD (Visual Question Answering in Radiology) is the first manually constructed dataset of medical visual question answering containing 3,515 clinically generated visual questions and answers about radiology images. The dataset includes questions created by clinical trainees on 315 radiology images from MedPix covering head, chest, and abdominal scans, designed to support AI development for medical image analysis and improve patient care.

VQA-Rad is a multimodal benchmark evaluating models on multimodal, image to text, healthcare, and vision tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for multimodal, best AI for image to text, best AI for healthcare and best AI for vision leaderboards.

Current leaders

MedGemma 4B IT from Google currently leads the VQA-Rad leaderboard with a score of 0.499 across 1 evaluated AI models.

1MedGemma 4B ITGoogle49.9%

FAQ

Common questions about the VQA-Rad benchmark and leaderboard.

What is the VQA-Rad benchmark?

VQA-RAD (Visual Question Answering in Radiology) is the first manually constructed dataset of medical visual question answering containing 3,515 clinically generated visual questions and answers about radiology images. The dataset includes questions created by clinical trainees on 315 radiology images from MedPix covering head, chest, and abdominal scans, designed to support AI development for medical image analysis and improve patient care.

What is the VQA-Rad leaderboard?

The VQA-Rad leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, MedGemma 4B IT by Google leads with a score of 0.499. The average score across all models is 0.499.

What is the highest VQA-Rad score?

The highest VQA-Rad score is 0.499, achieved by MedGemma 4B IT from Google.

How many models are evaluated on VQA-Rad?

1 models have been evaluated on the VQA-Rad benchmark, with 0 verified results and 1 self-reported results.

Where can I find the VQA-Rad paper?

The VQA-Rad paper is available at https://doi.org/10.1038/sdata.2018.251. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does VQA-Rad cover?

VQA-Rad is categorized under multimodal, image to text, healthcare, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on VQA-Rad?

MedGemma 4B IT by Google is the top-ranked open-source model on VQA-Rad, with a score of 0.499 (rank #1).

How recent are the VQA-Rad leaderboard results?

The VQA-Rad leaderboard was last updated in August 2026 and currently includes 1 evaluated models.