The AI arena is free today

Open Superagent

DeepPlanning

Progress Over Time

Interactive timeline showing model performance evolution on DeepPlanning

State-of-the-art frontier
Open
Proprietary

DeepPlanning Leaderboard

9 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
Notice missing or incorrect data?
About this benchmark

What is DeepPlanning?

DeepPlanning evaluates LLMs on complex multi-step planning tasks requiring long-horizon reasoning, goal decomposition, and strategic decision-making.

DeepPlanning is a text benchmark evaluating models on reasoning and agents tasks. LLM Stats tracks 9 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.6.

Compare leaders on the best AI for reasoning and best AI for agents leaderboards.

Current leaders

Qwen3.7-Plus from Alibaba Cloud / Qwen Team currently leads the DeepPlanning leaderboard with a score of 0.623 across 9 evaluated AI models.

1Qwen3.7-PlusAlibaba Cloud / Qwen Team62.3%
2Qwen3.6 PlusAlibaba Cloud / Qwen Team41.5%
3Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team34.3%

FAQ

Common questions about the DeepPlanning benchmark and leaderboard.

What is the DeepPlanning benchmark?

DeepPlanning evaluates LLMs on complex multi-step planning tasks requiring long-horizon reasoning, goal decomposition, and strategic decision-making.

What is the DeepPlanning leaderboard?

The DeepPlanning leaderboard ranks 9 AI models based on their performance on this benchmark. Currently, Qwen3.7-Plus by Alibaba Cloud / Qwen Team leads with a score of 0.623. The average score across all models is 0.299.

What is the highest DeepPlanning score?

The highest DeepPlanning score is 0.623, achieved by Qwen3.7-Plus from Alibaba Cloud / Qwen Team.

How many models are evaluated on DeepPlanning?

9 models have been evaluated on the DeepPlanning benchmark, with 0 verified results and 9 self-reported results.

What categories does DeepPlanning cover?

DeepPlanning is categorized under reasoning and agents. The benchmark evaluates text models.

What is the best open-source model on DeepPlanning?

Qwen3.5-397B-A17B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on DeepPlanning, with a score of 0.343 (rank #3).

How recent are the DeepPlanning leaderboard results?

The DeepPlanning leaderboard was last updated in August 2026 and currently includes 9 evaluated models.
DeepPlanning Leaderboard