The AI arena is free today

Open Superagent

DSBench-FullStack

Progress Over Time

Interactive timeline showing model performance evolution on DSBench-FullStack

State-of-the-art frontier
Open
Proprietary

DSBench-FullStack Leaderboard

2 models
ContextCostLicense
11.6T1.0M$0.43 / $0.87
2304B1.0M$0.09 / $0.18
Notice missing or incorrect data?
About this benchmark

What is DSBench-FullStack?

DSBench-FullStack is DeepSeek's internal full-stack development test set for evaluating coding agents on end-to-end software engineering tasks.

DSBench-FullStack is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

DeepSeek-V4-Pro-0813 from DeepSeek currently leads the DSBench-FullStack leaderboard with a score of 0.711 across 2 evaluated AI models.

1DeepSeek-V4-Pro-0813DeepSeek71.1%

FAQ

Common questions about the DSBench-FullStack benchmark and leaderboard.

What is the DSBench-FullStack benchmark?

DSBench-FullStack is DeepSeek's internal full-stack development test set for evaluating coding agents on end-to-end software engineering tasks.

What is the DSBench-FullStack leaderboard?

The DSBench-FullStack leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, DeepSeek-V4-Pro-0813 by DeepSeek leads with a score of 0.711. The average score across all models is 0.699.

What is the highest DSBench-FullStack score?

The highest DSBench-FullStack score is 0.711, achieved by DeepSeek-V4-Pro-0813 from DeepSeek.

How many models are evaluated on DSBench-FullStack?

2 models have been evaluated on the DSBench-FullStack benchmark, with 0 verified results and 2 self-reported results.

What categories does DSBench-FullStack cover?

DSBench-FullStack is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on DSBench-FullStack?

DeepSeek-V4-Pro-0813 by DeepSeek is the top-ranked open-source model on DSBench-FullStack, with a score of 0.711 (rank #1).

Which model offers the best value on DSBench-FullStack?

Among models scoring within 10% of the leader, DeepSeek-V4-Flash-0731 from DeepSeek is the cheapest, at $0.09 per million input tokens with a score of 0.687.

How recent are the DSBench-FullStack leaderboard results?

The DSBench-FullStack leaderboard was last updated in August 2026 and currently includes 2 evaluated models.