MedXpertQA
Progress Over Time
Interactive timeline showing model performance evolution on MedXpertQA
MedXpertQA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Meta | — | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | 122B | 262K | $0.29 / $2.40 | ||
| 3 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.26 / $2.60 | ||
| 4 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.14 / $1.00 | ||
| 5 | Google | 31B | 262K | $0.09 / $0.34 | ||
| 6 | Google | 25B | 262K | $0.07 / $0.34 | ||
| 7 | Google | 25B | — | — | ||
| 8 | Google | 12B | — | — | ||
| 9 | Microsoft | 1.0T | — | — | ||
| 10 | Google | 8B | 131K | $0.02 / $0.10 | ||
| 11 | Google | 5B | — | — | ||
| 12 | Google | 4B | — | — |
What is MedXpertQA?
A comprehensive benchmark to evaluate expert-level medical knowledge and advanced reasoning, featuring 4,460 questions spanning 17 specialties and 11 body systems. Includes both text-only and multimodal subsets with expert-level exam questions incorporating diverse medical images and rich clinical information.
MedXpertQA is a multimodal benchmark evaluating models on multimodal, reasoning, healthcare, and vision tasks. LLM Stats tracks 12 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.
Compare leaders on the best AI for multimodal, best AI for reasoning, best AI for healthcare and best AI for vision leaderboards.
Current leaders
Muse Spark from Meta currently leads the MedXpertQA leaderboard with a score of 0.784 across 12 evaluated AI models.
FAQ
Common questions about the MedXpertQA benchmark and leaderboard.