SlakeVQA
Progress Over Time
Interactive timeline showing model performance evolution on SlakeVQA
SlakeVQA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 3 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 4 | Google | 4B | — | — |
What is SlakeVQA?
A semantically-labeled knowledge-enhanced dataset for medical visual question answering. Contains 642 radiology images (CT scans, MRI scans, X-rays) covering five body parts and 14,028 bilingual English-Chinese question-answer pairs annotated by experienced physicians. Features comprehensive semantic labels and a structural medical knowledge base with both vision-only and knowledge-based questions requiring external medical knowledge reasoning.
SlakeVQA is a multimodal benchmark evaluating models on multimodal, reasoning, image to text, healthcare, and vision tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for multimodal, best AI for reasoning, best AI for image to text, best AI for healthcare and best AI for vision leaderboards.
Current leaders
Qwen3.5-122B-A10B from Alibaba Cloud / Qwen Team currently leads the SlakeVQA leaderboard with a score of 0.816 across 4 evaluated AI models.
Source paper
- Title
- SLAKE: A Semantically-Labeled Knowledge-Enhanced Dataset for Medical Visual Question Answering
- Authors
- Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, and 2 others
- Published
- arXiv
- 2102.09542
Abstract
Medical visual question answering (Med-VQA) has tremendous potential in healthcare. However, the development of this technology is hindered by the lacking of publicly-available and high-quality labeled datasets for training and evaluation. In this paper, we present a large bilingual dataset, SLAKE, with comprehensive semantic labels annotated by experienced physicians and a new structural medical knowledge base for Med-VQA. Besides, SLAKE includes richer modalities and covers more human body parts than the currently available dataset. We show that SLAKE can be used to facilitate the development and evaluation of Med-VQA systems. The dataset can be downloaded from http://www.med-vqa.com/slake.
FAQ
Common questions about the SlakeVQA benchmark and leaderboard.