The AI arena is free today

Open Superagent

Workspace Bench

Progress Over Time

Interactive timeline showing model performance evolution on Workspace Bench

State-of-the-art frontier
Open
Proprietary

Workspace Bench Leaderboard

3 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
3
ByteDance
ByteDance
Notice missing or incorrect data?
About this benchmark

What is Workspace Bench?

Workspace Bench evaluates AI agents on high-economic-value workplace tasks that span multi-step planning, file processing, and tool use across realistic office and productivity workflows.

Workspace Bench is a text benchmark evaluating models on reasoning, general, and agents tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.

Compare leaders on the best AI for reasoning, best AI for general and best AI for agents leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the Workspace Bench leaderboard with a score of 0.677 across 3 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team67.7%
2Seed 2.1 TurboByteDance54.7%
3Seed 2.1 ProByteDance53.0%

FAQ

Common questions about the Workspace Bench benchmark and leaderboard.

What is the Workspace Bench benchmark?

Workspace Bench evaluates AI agents on high-economic-value workplace tasks that span multi-step planning, file processing, and tool use across realistic office and productivity workflows.

What is the Workspace Bench leaderboard?

The Workspace Bench leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.677. The average score across all models is 0.585.

What is the highest Workspace Bench score?

The highest Workspace Bench score is 0.677, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on Workspace Bench?

3 models have been evaluated on the Workspace Bench benchmark, with 0 verified results and 3 self-reported results.

What categories does Workspace Bench cover?

Workspace Bench is categorized under reasoning, general, and agents. The benchmark evaluates text models.

What is the best open-source model on Workspace Bench?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on Workspace Bench, with a score of 0.677 (rank #1).

Which model offers the best value on Workspace Bench?

Among models scoring within 10% of the leader, Qwen3.8 Max from Alibaba Cloud / Qwen Team is the cheapest, at $1.65 per million input tokens with a score of 0.677.

How recent are the Workspace Bench leaderboard results?

The Workspace Bench leaderboard was last updated in August 2026 and currently includes 3 evaluated models.