The AI arena is free today

Open Superagent

WideSearch

Progress Over Time

Interactive timeline showing model performance evolution on WideSearch

State-of-the-art frontier
Open
Proprietary

WideSearch Leaderboard

10 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
3
Moonshot AI
Moonshot AI
1.0T
4
Tencent
Tencent
295B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is WideSearch?

WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.

WideSearch is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 10 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the WideSearch leaderboard with a score of 0.819 across 10 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team81.9%
2Kimi K2.6Moonshot AI80.8%
3Kimi K2.5Moonshot AI79.0%

FAQ

Common questions about the WideSearch benchmark and leaderboard.

What is the WideSearch benchmark?

WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.

What is the WideSearch leaderboard?

The WideSearch leaderboard ranks 10 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.819. The average score across all models is 0.705.

What is the highest WideSearch score?

The highest WideSearch score is 0.819, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on WideSearch?

10 models have been evaluated on the WideSearch benchmark, with 0 verified results and 10 self-reported results.

What categories does WideSearch cover?

WideSearch is categorized under reasoning, search, and agents. The benchmark evaluates text models.

What is the best open-source model on WideSearch?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on WideSearch, with a score of 0.819 (rank #1).

Which model offers the best value on WideSearch?

Among models scoring within 10% of the leader, Qwen3.6 Plus from Alibaba Cloud / Qwen Team is the cheapest, at $0.50 per million input tokens with a score of 0.743.

How recent are the WideSearch leaderboard results?

The WideSearch leaderboard was last updated in August 2026 and currently includes 10 evaluated models.