DermMCQA
Progress Over Time
Interactive timeline showing model performance evolution on DermMCQA
DermMCQA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | 4B | — | — |
What is DermMCQA?
Dermatology multiple choice question assessment benchmark for evaluating medical knowledge and diagnostic reasoning in dermatological conditions and treatments.
DermMCQA is a text benchmark evaluating models on healthcare tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for healthcare leaderboards.
Current leaders
MedGemma 4B IT from Google currently leads the DermMCQA leaderboard with a score of 0.718 across 1 evaluated AI models.
Source paper
- Title
- Towards Reliable Dermatology Evaluation Benchmarks
- Authors
- Fabian Gröger, Simone Lionetti, Philippe Gottfrois, Alvaro Gonzalez-Jimenez, and 5 others
- Published
- arXiv
- 2309.06961
Abstract
Benchmark datasets for digital dermatology unwittingly contain inaccuracies that reduce trust in model performance estimates. We propose a resource-efficient data-cleaning protocol to identify issues that escaped previous curation. The protocol leverages an existing algorithmic cleaning strategy and is followed by a confirmation process terminated by an intuitive stopping criterion. Based on confirmation by multiple dermatologists, we remove irrelevant samples and near duplicates and estimate the percentage of label errors in six dermatology image datasets for model evaluation promoted by the International Skin Imaging Collaboration. Along with this paper, we publish revised file lists for each dataset which should be used for model evaluation. Our work paves the way for more trustworthy performance assessment in digital dermatology.
FAQ
Common questions about the DermMCQA benchmark and leaderboard.