The AI arena is free today

Open Playground

DSBench-Hard

Progress Over Time

Interactive timeline showing model performance evolution on DSBench-Hard

State-of-the-art frontier
Open
Proprietary

DSBench-Hard Leaderboard

1 models
ContextCostLicense
1304B1.0M$0.09 / $0.18
Notice missing or incorrect data?
About this benchmark

What is DSBench-Hard?

DSBench-Hard is DeepSeek's internal test set of difficult coding-agent problems.

DSBench-Hard is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

DeepSeek-V4-Flash-0731 from DeepSeek currently leads the DSBench-Hard leaderboard with a score of 0.596 across 1 evaluated AI models.

FAQ

Common questions about the DSBench-Hard benchmark and leaderboard.

What is the DSBench-Hard benchmark?

DSBench-Hard is DeepSeek's internal test set of difficult coding-agent problems.

What is the DSBench-Hard leaderboard?

The DSBench-Hard leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, DeepSeek-V4-Flash-0731 by DeepSeek leads with a score of 0.596. The average score across all models is 0.596.

What is the highest DSBench-Hard score?

The highest DSBench-Hard score is 0.596, achieved by DeepSeek-V4-Flash-0731 from DeepSeek.

How many models are evaluated on DSBench-Hard?

1 models have been evaluated on the DSBench-Hard benchmark, with 0 verified results and 1 self-reported results.

What categories does DSBench-Hard cover?

DSBench-Hard is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on DSBench-Hard?

DeepSeek-V4-Flash-0731 by DeepSeek is the top-ranked open-source model on DSBench-Hard, with a score of 0.596 (rank #1).

Which model offers the best value on DSBench-Hard?

Among models scoring within 10% of the leader, DeepSeek-V4-Flash-0731 from DeepSeek is the cheapest, at $0.09 per million input tokens with a score of 0.596.

How recent are the DSBench-Hard leaderboard results?

The DSBench-Hard leaderboard was last updated in August 2026 and currently includes 1 evaluated models.