The AI arena is free today

Open Superagent

OneMillion Bench

Progress Over Time

Interactive timeline showing model performance evolution on OneMillion Bench

State-of-the-art frontier
Open
Proprietary

OneMillion Bench Leaderboard

3 models
ContextCostLicense
1
ByteDance
ByteDance
2
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
Notice missing or incorrect data?
About this benchmark

What is OneMillion Bench?

OneMillion Bench evaluates AI agents on high-economic-value tasks that require sustained, reliable execution across long-horizon real-world workflows.

OneMillion Bench is a text benchmark evaluating models on long context, reasoning, and agents tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.

Compare leaders on the best AI for long context, best AI for reasoning and best AI for agents leaderboards.

Current leaders

Seed 2.1 Pro from ByteDance currently leads the OneMillion Bench leaderboard with a score of 0.688 across 3 evaluated AI models.

1Seed 2.1 ProByteDance68.8%
2Seed 2.1 TurboByteDance66.6%
3Qwen3.8 MaxAlibaba Cloud / Qwen Team52.5%

FAQ

Common questions about the OneMillion Bench benchmark and leaderboard.

What is the OneMillion Bench benchmark?

OneMillion Bench evaluates AI agents on high-economic-value tasks that require sustained, reliable execution across long-horizon real-world workflows.

What is the OneMillion Bench leaderboard?

The OneMillion Bench leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Seed 2.1 Pro by ByteDance leads with a score of 0.688. The average score across all models is 0.626.

What is the highest OneMillion Bench score?

The highest OneMillion Bench score is 0.688, achieved by Seed 2.1 Pro from ByteDance.

How many models are evaluated on OneMillion Bench?

3 models have been evaluated on the OneMillion Bench benchmark, with 0 verified results and 3 self-reported results.

What categories does OneMillion Bench cover?

OneMillion Bench is categorized under long context, reasoning, and agents. The benchmark evaluates text models.

What is the best open-source model on OneMillion Bench?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on OneMillion Bench, with a score of 0.525 (rank #3).

How recent are the OneMillion Bench leaderboard results?

The OneMillion Bench leaderboard was last updated in August 2026 and currently includes 3 evaluated models.