The AI arena is free today

Open Superagent

FrontierScience Research

Progress Over Time

Interactive timeline showing model performance evolution on FrontierScience Research

State-of-the-art frontier
Open
Proprietary

FrontierScience Research Leaderboard

4 models
ContextCostLicense
1
2
3
ByteDance
ByteDance
4
Tencent
Tencent
295B
Notice missing or incorrect data?
About this benchmark

What is FrontierScience Research?

FrontierScience Research is a benchmark evaluating AI models on cutting-edge scientific research questions requiring deep domain expertise, multi-step reasoning, and synthesis of complex scientific concepts across disciplines.

FrontierScience Research is a text benchmark evaluating models on reasoning and science tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.

Compare leaders on the best AI for reasoning and best AI for science leaderboards.

Current leaders

Muse Spark from Meta currently leads the FrontierScience Research leaderboard with a score of 0.383 across 4 evaluated AI models.

1Muse SparkMeta38.3%
2Seed 2.1 TurboByteDance33.3%
3Seed 2.1 ProByteDance28.3%
OSSHy3#4 open-weight21.3%

FAQ

Common questions about the FrontierScience Research benchmark and leaderboard.

What is the FrontierScience Research benchmark?

FrontierScience Research is a benchmark evaluating AI models on cutting-edge scientific research questions requiring deep domain expertise, multi-step reasoning, and synthesis of complex scientific concepts across disciplines.

What is the FrontierScience Research leaderboard?

The FrontierScience Research leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, Muse Spark by Meta leads with a score of 0.383. The average score across all models is 0.303.

What is the highest FrontierScience Research score?

The highest FrontierScience Research score is 0.383, achieved by Muse Spark from Meta.

How many models are evaluated on FrontierScience Research?

4 models have been evaluated on the FrontierScience Research benchmark, with 0 verified results and 4 self-reported results.

What categories does FrontierScience Research cover?

FrontierScience Research is categorized under reasoning and science. The benchmark evaluates text models.

What is the best open-source model on FrontierScience Research?

Hy3 by Tencent is the top-ranked open-source model on FrontierScience Research, with a score of 0.213 (rank #4).

How recent are the FrontierScience Research leaderboard results?

The FrontierScience Research leaderboard was last updated in August 2026 and currently includes 4 evaluated models.
FrontierScience Research Leaderboard