The AI arena is free today

Open Superagent

DeepSearchQA

Progress Over Time

Interactive timeline showing model performance evolution on DeepSearchQA

State-of-the-art frontier
Open
Proprietary

DeepSearchQA Leaderboard

9 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
21.0M$5.00 / $25.00
31.0M$5.00 / $25.00
4
Tencent
Tencent
295B
51.0T
6
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
7
Moonshot AI
Moonshot AI
1.0T
8
930B
Notice missing or incorrect data?
About this benchmark

What is DeepSearchQA?

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.

DeepSearchQA is a text benchmark evaluating models on reasoning, search, and agents tasks. LLM Stats tracks 9 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for search and best AI for agents leaderboards.

Current leaders

Kimi K3 from Moonshot AI currently leads the DeepSearchQA leaderboard with a score of 0.950 across 9 evaluated AI models.

1Kimi K3Moonshot AI95.0%
2Claude Opus 4.8Anthropic93.1%
3Claude Opus 4.6Anthropic91.3%

FAQ

Common questions about the DeepSearchQA benchmark and leaderboard.

What is the DeepSearchQA benchmark?

DeepSearchQA is a benchmark for evaluating deep search and question-answering capabilities, testing models' ability to perform multi-hop reasoning and information retrieval across complex knowledge domains.

What is the DeepSearchQA leaderboard?

The DeepSearchQA leaderboard ranks 9 AI models based on their performance on this benchmark. Currently, Kimi K3 by Moonshot AI leads with a score of 0.950. The average score across all models is 0.852.

What is the highest DeepSearchQA score?

The highest DeepSearchQA score is 0.950, achieved by Kimi K3 from Moonshot AI.

How many models are evaluated on DeepSearchQA?

9 models have been evaluated on the DeepSearchQA benchmark, with 0 verified results and 9 self-reported results.

What categories does DeepSearchQA cover?

DeepSearchQA is categorized under reasoning, search, and agents. The benchmark evaluates text models.

What is the best open-source model on DeepSearchQA?

Kimi K3 by Moonshot AI is the top-ranked open-source model on DeepSearchQA, with a score of 0.950 (rank #1).

Which model offers the best value on DeepSearchQA?

Among models scoring within 10% of the leader, Kimi K3 from Moonshot AI is the cheapest, at $3.00 per million input tokens with a score of 0.950.

How recent are the DeepSearchQA leaderboard results?

The DeepSearchQA leaderboard was last updated in August 2026 and currently includes 9 evaluated models.