BioMysteryBench (Human Difficult)
Progress Over Time
Interactive timeline showing model performance evolution on BioMysteryBench (Human Difficult)
BioMysteryBench (Human Difficult) Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Anthropic | — | 1.0M | $2.00 / $10.00 |
What is BioMysteryBench (Human Difficult)?
BioMysteryBench Human Difficult subset: problems that human experts did not solve, separate from the Human Solvable subset.
BioMysteryBench (Human Difficult) is a text benchmark evaluating models on reasoning, science, and biology tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.4.
Compare leaders on the best AI for reasoning, best AI for science and best AI for biology leaderboards.
Current leaders
Claude Sonnet 5.5 from Anthropic currently leads the BioMysteryBench (Human Difficult) leaderboard with a score of 0.447 across 1 evaluated AI models.
FAQ
Common questions about the BioMysteryBench (Human Difficult) benchmark and leaderboard.