HealthBench Hard
Progress Over Time
Interactive timeline showing model performance evolution on HealthBench Hard
HealthBench Hard Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Meta | — | — | — | ||
| 2 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 3 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 4 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 5 | OpenAI | 117B | 131K | $0.04 / $0.17 | ||
| 6 | OpenAI | — | 128K | $1.75 / $14.00 | ||
| 7 | OpenAI | — | 400K | $5.00 / $30.00 | ||
| 8 | ByteDance | — | 256K | $0.10 / $0.40 | ||
| 9 | OpenAI | 21B | 131K | $0.03 / $0.14 | ||
| 10 | OpenAI | — | — | — |
What is HealthBench Hard?
A challenging variation of HealthBench that evaluates large language models' performance and safety in healthcare through 5,000 multi-turn conversations with particularly rigorous evaluation criteria validated by 262 physicians from 60 countries
HealthBench Hard is a text benchmark evaluating models on healthcare tasks. LLM Stats tracks 10 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.4.
Compare leaders on the best AI for healthcare leaderboards.
Current leaders
Muse Spark from Meta currently leads the HealthBench Hard leaderboard with a score of 0.428 across 10 evaluated AI models.
FAQ
Common questions about the HealthBench Hard benchmark and leaderboard.