GeneBench
Progress Over Time
Interactive timeline showing model performance evolution on GeneBench
GeneBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | — | — | ||
| 2 | OpenAI | — | 1.1M | $5.00 / $30.00 |
Sub-benchmarks
What is GeneBench?
GeneBench is an evaluation focused on multi-stage scientific data analysis in genetics and quantitative biology. Tasks require reasoning about ambiguous or noisy data with minimal supervisory guidance, addressing realistic obstacles such as hidden confounders or QC failures, and correctly implementing and interpreting modern statistical methods.
GeneBench is a text benchmark evaluating models on reasoning, science, and agents tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.3.
Compare leaders on the best AI for reasoning, best AI for science and best AI for agents leaderboards.
Current leaders
GPT-5.5 Pro from OpenAI currently leads the GeneBench leaderboard with a score of 0.332 across 2 evaluated AI models.
FAQ
Common questions about the GeneBench benchmark and leaderboard.