The AI arena is free today

Open Superagent

WideSearch

Progress Over Time

Interactive timeline showing model performance evolution on WideSearch

State-of-the-art frontier
Open
Proprietary

WideSearch Leaderboard

11 models
ContextCostLicense
1
Tencent
Tencent
770B
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
3
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
4
Moonshot AI
Moonshot AI
1.0T
5
Tencent
Tencent
295B
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is WideSearch?

WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.

WideSearch is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.

Current leaders

Hy4 preview from Tencent currently leads the WideSearch leaderboard with a score of 0.839 across 11 evaluated AI models.

1Hy4 previewTencent83.9%
2Qwen3.8 MaxAlibaba Cloud / Qwen Team81.9%
3Kimi K2.6Moonshot AI80.8%

FAQ

Common questions about the WideSearch benchmark and leaderboard.

What is the WideSearch benchmark?

WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.

What is the WideSearch leaderboard?

The WideSearch leaderboard ranks 11 AI models based on their performance on this benchmark. Currently, Hy4 preview by Tencent leads with a score of 0.839. The average score across all models is 0.717.

What is the highest WideSearch score?

The highest WideSearch score is 0.839, achieved by Hy4 preview from Tencent.

How many models are evaluated on WideSearch?

11 models have been evaluated on the WideSearch benchmark, with 0 verified results and 11 self-reported results.

What categories does WideSearch cover?

WideSearch is categorized under reasoning, search, and agents. The benchmark evaluates text models.

What is the best open-source model on WideSearch?

Hy4 preview by Tencent is the top-ranked open-source model on WideSearch, with a score of 0.839 (rank #1).

Which model offers the best value on WideSearch?

Among models scoring within 10% of the leader, Kimi K2.6 from Moonshot AI is the cheapest, at $0.75 per million input tokens with a score of 0.808.

How recent are the WideSearch leaderboard results?

The WideSearch leaderboard was last updated in September 2026 and currently includes 11 evaluated models.