WideSearch
Progress Over Time
Interactive timeline showing model performance evolution on WideSearch
WideSearch Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 2 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 3 | Moonshot AI | 1.0T | — | — | ||
| 4 | Tencent | 295B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 6 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 7 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 8 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 9 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 10 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is WideSearch?
WideSearch is an agentic search benchmark that evaluates models' ability to perform broad, parallel search operations across multiple sources. It tests wide-coverage information retrieval and synthesis capabilities.
WideSearch is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 10 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.
Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.
Current leaders
Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the WideSearch leaderboard with a score of 0.819 across 10 evaluated AI models.
FAQ
Common questions about the WideSearch benchmark and leaderboard.