The AI arena is free today

Open Superagent

DeepSearchQA

Progress Over Time

Interactive timeline showing model performance evolution on DeepSearchQA

State-of-the-art frontier
Open
Proprietary

DeepSearchQA Leaderboard

10 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
21.0M$5.00 / $25.00
31.0M$5.00 / $25.00
4
Tencent
Tencent
295B
51.0M$0.10 / $0.20
61.0T
7
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
8
Moonshot AI
Moonshot AI
1.0T
9
1030B
Notice missing or incorrect data?
About this benchmark

What is DeepSearchQA?

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.

DeepSearchQA is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 10 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.

Current leaders

Kimi K3 from Moonshot AI currently leads the DeepSearchQA leaderboard with a score of 0.950 across 10 evaluated AI models.

1Kimi K3Moonshot AI95.0%
2Claude Opus 4.8Anthropic93.1%
3Claude Opus 4.6Anthropic91.3%
OSSHy3#4 open-weight91.0%

FAQ

Common questions about the DeepSearchQA benchmark and leaderboard.

What is the DeepSearchQA benchmark?

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.

What is the DeepSearchQA leaderboard?

The DeepSearchQA leaderboard ranks 10 AI models based on their performance on this benchmark. Currently, Kimi K3 by Moonshot AI leads with a score of 0.950. The average score across all models is 0.856.

What is the highest DeepSearchQA score?

The highest DeepSearchQA score is 0.950, achieved by Kimi K3 from Moonshot AI.

How many models are evaluated on DeepSearchQA?

10 models have been evaluated on the DeepSearchQA benchmark, with 0 verified results and 10 self-reported results.

What categories does DeepSearchQA cover?

DeepSearchQA is categorized under reasoning, search, and agents. The benchmark evaluates text models.

What is the best open-source model on DeepSearchQA?

Hy3 by Tencent is the top-ranked open-source model on DeepSearchQA, with a score of 0.910 (rank #4).

Which model offers the best value on DeepSearchQA?

Among models scoring within 10% of the leader, Muse Spark 1.3 from Meta is the cheapest, at $0.10 per million input tokens with a score of 0.894.

How recent are the DeepSearchQA leaderboard results?

The DeepSearchQA leaderboard was last updated in September 2026 and currently includes 10 evaluated models.