The AI arena is free today

Open Playground

DSBench-FullStack

Progress Over Time

Interactive timeline showing model performance evolution on DSBench-FullStack

State-of-the-art frontier
Open
Proprietary

DSBench-FullStack Leaderboard

1 models
ContextCostLicense
1304B1.0M$0.09 / $0.18
Notice missing or incorrect data?
About this benchmark

What is DSBench-FullStack?

DSBench-FullStack is DeepSeek's internal full-stack development test set for evaluating coding agents on end-to-end software engineering tasks.

DSBench-FullStack is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

DeepSeek-V4-Flash-0731 from DeepSeek currently leads the DSBench-FullStack leaderboard with a score of 0.687 across 1 evaluated AI models.

FAQ

Common questions about the DSBench-FullStack benchmark and leaderboard.

What is the DSBench-FullStack benchmark?

DSBench-FullStack is DeepSeek's internal full-stack development test set for evaluating coding agents on end-to-end software engineering tasks.

What is the DSBench-FullStack leaderboard?

The DSBench-FullStack leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, DeepSeek-V4-Flash-0731 by DeepSeek leads with a score of 0.687. The average score across all models is 0.687.

What is the highest DSBench-FullStack score?

The highest DSBench-FullStack score is 0.687, achieved by DeepSeek-V4-Flash-0731 from DeepSeek.

How many models are evaluated on DSBench-FullStack?

1 models have been evaluated on the DSBench-FullStack benchmark, with 0 verified results and 1 self-reported results.

What categories does DSBench-FullStack cover?

DSBench-FullStack is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on DSBench-FullStack?

DeepSeek-V4-Flash-0731 by DeepSeek is the top-ranked open-source model on DSBench-FullStack, with a score of 0.687 (rank #1).

Which model offers the best value on DSBench-FullStack?

Among models scoring within 10% of the leader, DeepSeek-V4-Flash-0731 from DeepSeek is the cheapest, at $0.09 per million input tokens with a score of 0.687.

How recent are the DSBench-FullStack leaderboard results?

The DSBench-FullStack leaderboard was last updated in August 2026 and currently includes 1 evaluated models.