The AI arena is free today

Open Superagent

CoWorkBench

Progress Over Time

Interactive timeline showing model performance evolution on CoWorkBench

State-of-the-art frontier
Open
Proprietary

CoWorkBench Leaderboard

5 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
Notice missing or incorrect data?
About this benchmark

What is CoWorkBench?

CoWorkBench is Qwen's internal cowork benchmark for evaluating long-horizon office and productivity agent tasks across domains such as computer science, finance, law, and medicine.

CoWorkBench is a text benchmark evaluating models on productivity, reasoning, and agents tasks. LLM Stats tracks 5 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for productivity, best AI for reasoning and best AI for agents leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the CoWorkBench leaderboard with a score of 0.748 across 5 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team74.8%
2Qwen3.8-Flash-NextAlibaba Cloud / Qwen Team73.9%
3Qwen3.8-27BAlibaba Cloud / Qwen Team70.7%

FAQ

Common questions about the CoWorkBench benchmark and leaderboard.

What is the CoWorkBench benchmark?

CoWorkBench is Qwen's internal cowork benchmark for evaluating long-horizon office and productivity agent tasks across domains such as computer science, finance, law, and medicine.

What is the CoWorkBench leaderboard?

The CoWorkBench leaderboard ranks 5 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.748. The average score across all models is 0.703.

What is the highest CoWorkBench score?

The highest CoWorkBench score is 0.748, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on CoWorkBench?

5 models have been evaluated on the CoWorkBench benchmark, with 0 verified results and 5 self-reported results.

What categories does CoWorkBench cover?

CoWorkBench is categorized under productivity, reasoning, and agents. The benchmark evaluates text models.

What is the best open-source model on CoWorkBench?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on CoWorkBench, with a score of 0.748 (rank #1).

Which model offers the best value on CoWorkBench?

Among models scoring within 10% of the leader, Qwen3.8 Max from Alibaba Cloud / Qwen Team is the cheapest, at $1.65 per million input tokens with a score of 0.748.

How recent are the CoWorkBench leaderboard results?

The CoWorkBench leaderboard was last updated in August 2026 and currently includes 5 evaluated models.