DeepSearchQA
Progress Over Time
Interactive timeline showing model performance evolution on DeepSearchQA
DeepSearchQA Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 3 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 4 | Tencent | 295B | — | — | ||
| 5 | Xiaomi | 1.0T | — | — | ||
| 6 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 7 | Moonshot AI | 1.0T | — | — | ||
| 8 | Meta | — | — | — | ||
| 9 | Meta | 30B | — | — |
What is DeepSearchQA?
DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.
DeepSearchQA is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 9 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.
Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the DeepSearchQA leaderboard with a score of 0.950 across 9 evaluated AI models.
FAQ
Common questions about the DeepSearchQA benchmark and leaderboard.