DSBench-FullStack
Progress Over Time
Interactive timeline showing model performance evolution on DSBench-FullStack
DSBench-FullStack Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | DeepSeek | 304B | 1.0M | $0.09 / $0.18 |
What is DSBench-FullStack?
DSBench-FullStack is DeepSeek's internal full-stack development test set for evaluating coding agents on end-to-end software engineering tasks.
DSBench-FullStack is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
DeepSeek-V4-Flash-0731 from DeepSeek currently leads the DSBench-FullStack leaderboard with a score of 0.687 across 1 evaluated AI models.
FAQ
Common questions about the DSBench-FullStack benchmark and leaderboard.