The AI arena is free today

Open Superagent

PaperBench

Progress Over Time

Interactive timeline showing model performance evolution on PaperBench

State-of-the-art frontier
Open
Proprietary

PaperBench Leaderboard

3 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
Moonshot AI
Moonshot AI
1.0T
3
MiniMax
MiniMax
428B1.0M$0.30 / $1.20
Notice missing or incorrect data?
About this benchmark

What is PaperBench?

PaperBench is a benchmark for evaluating AI agents on their ability to replicate research papers. It tests models on complex, multi-step workflows involving code implementation, experimentation, and reproducing scientific results from academic publications.

PaperBench is a text benchmark evaluating models on reasoning, agents, and code tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for agents and best AI for code leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the PaperBench leaderboard with a score of 0.930 across 3 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team93.0%
2Kimi K2.5Moonshot AI63.5%
3MiniMax M3MiniMax52.6%

FAQ

Common questions about the PaperBench benchmark and leaderboard.

What is the PaperBench benchmark?

PaperBench is a benchmark for evaluating AI agents on their ability to replicate research papers. It tests models on complex, multi-step workflows involving code implementation, experimentation, and reproducing scientific results from academic publications.

What is the PaperBench leaderboard?

The PaperBench leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.930. The average score across all models is 0.697.

What is the highest PaperBench score?

The highest PaperBench score is 0.930, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on PaperBench?

3 models have been evaluated on the PaperBench benchmark, with 0 verified results and 3 self-reported results.

What categories does PaperBench cover?

PaperBench is categorized under reasoning, agents, and code. The benchmark evaluates text models.

What is the best open-source model on PaperBench?

Kimi K2.5 by Moonshot AI is the top-ranked open-source model on PaperBench, with a score of 0.635 (rank #2).

Which model offers the best value on PaperBench?

Among models scoring within 10% of the leader, Qwen3.8 Max from Alibaba Cloud / Qwen Team is the cheapest, at $1.65 per million input tokens with a score of 0.930.

How recent are the PaperBench leaderboard results?

The PaperBench leaderboard was last updated in September 2026 and currently includes 3 evaluated models.