The AI arena is free today

Open Superagent

LABBench2

Progress Over Time

Interactive timeline showing model performance evolution on LABBench2

State-of-the-art frontier
Open
Proprietary

LABBench2 Leaderboard

1 models
ContextCostLicense
11.0M$0.75 / $3.75
Notice missing or incorrect data?
About this benchmark

What is LABBench2?

LABBench2 evaluates models on real-world biology research tasks.

LABBench2 is a text benchmark evaluating models on reasoning, science, agents, and biology tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for science, best AI for agents and best AI for biology leaderboards.

Current leaders

Gemini 3.7 Flash from Google currently leads the LABBench2 leaderboard with a score of 0.821 across 1 evaluated AI models.

1Gemini 3.7 FlashGoogle82.1%

FAQ

Common questions about the LABBench2 benchmark and leaderboard.

What is the LABBench2 benchmark?

LABBench2 evaluates models on real-world biology research tasks.

What is the LABBench2 leaderboard?

The LABBench2 leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Gemini 3.7 Flash by Google leads with a score of 0.821. The average score across all models is 0.821.

What is the highest LABBench2 score?

The highest LABBench2 score is 0.821, achieved by Gemini 3.7 Flash from Google.

How many models are evaluated on LABBench2?

1 models have been evaluated on the LABBench2 benchmark, with 0 verified results and 1 self-reported results.

What categories does LABBench2 cover?

LABBench2 is categorized under reasoning, science, agents, and biology. The benchmark evaluates text models.

Which model offers the best value on LABBench2?

Among models scoring within 10% of the leader, Gemini 3.7 Flash from Google is the cheapest, at $0.75 per million input tokens with a score of 0.821.

How recent are the LABBench2 leaderboard results?

The LABBench2 leaderboard was last updated in August 2026 and currently includes 1 evaluated models.