LABBench2
Progress Over Time
Interactive timeline showing model performance evolution on LABBench2
State-of-the-art frontier
Open
Proprietary
LABBench2 Leaderboard
1 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | 1.0M | $0.75 / $3.75 |
Notice missing or incorrect data?
What is LABBench2?
LABBench2 evaluates models on real-world biology research tasks.
LABBench2 is a text benchmark evaluating models on reasoning, science, agents, and biology tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for science, best AI for agents and best AI for biology leaderboards.
Current leaders
Gemini 3.7 Flash from Google currently leads the LABBench2 leaderboard with a score of 0.821 across 1 evaluated AI models.
FAQ
Common questions about the LABBench2 benchmark and leaderboard.