BioLP-Bench
Progress Over Time
Interactive timeline showing model performance evolution on BioLP-Bench
BioLP-Bench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | — | — | — |
What is BioLP-Bench?
BioLP-Bench is a model-graded evaluation measuring ability to find and correct mistakes in common biological laboratory protocols. It evaluates dual-use biological knowledge relevant to bioweapons development.
BioLP-Bench is a text benchmark evaluating models on safety, healthcare, and biology tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.4.
Compare leaders on the best AI for safety, best AI for healthcare and best AI for biology leaderboards.
Current leaders
Grok-4.1 Thinking from xAI currently leads the BioLP-Bench leaderboard with a score of 0.370 across 1 evaluated AI models.
FAQ
Common questions about the BioLP-Bench benchmark and leaderboard.