SkillsBench
Progress Over Time
Interactive timeline showing model performance evolution on SkillsBench
SkillsBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 3 | Tencent | 295B | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.60 / $3.60 | ||
| 6 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 7 | Meta | 30B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is SkillsBench?
SkillsBench evaluates coding agents on self-contained programming tasks, measuring practical engineering skills across diverse software development scenarios.
SkillsBench is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 8 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the SkillsBench leaderboard with a score of 0.702 across 8 evaluated AI models.
FAQ
Common questions about the SkillsBench benchmark and leaderboard.