WideSearch
Progress Over Time
Interactive timeline showing model performance evolution on WideSearch
WideSearch Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Hy4 previewNew Tencent | 770B | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 3 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 4 | Moonshot AI | 1.0T | — | — | ||
| 5 | Tencent | 295B | — | — | ||
| 6 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 7 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 9 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 10 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 11 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is WideSearch?
WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.
WideSearch is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.
Current leaders
Hy4 preview from Tencent currently leads the WideSearch leaderboard with a score of 0.839 across 11 evaluated AI models.
FAQ
Common questions about the WideSearch benchmark and leaderboard.