The AI arena is free today

Open Superagent

SkillsBench

Progress Over Time

Interactive timeline showing model performance evolution on SkillsBench

State-of-the-art frontier
Open
Proprietary

SkillsBench Leaderboard

8 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
3
Tencent
Tencent
295B
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.60 / $3.60
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
730B
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is SkillsBench?

SkillsBench evaluates coding agents on self-contained programming tasks, measuring practical engineering skills across diverse software development scenarios.

SkillsBench is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 8 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Qwen3.8 Max from Alibaba Cloud / Qwen Team currently leads the SkillsBench leaderboard with a score of 0.702 across 8 evaluated AI models.

1Qwen3.8 MaxAlibaba Cloud / Qwen Team70.2%
2Qwen3.7 MaxAlibaba Cloud / Qwen Team59.2%
3Hy3Tencent55.3%

FAQ

Common questions about the SkillsBench benchmark and leaderboard.

What is the SkillsBench benchmark?

SkillsBench evaluates coding agents on self-contained programming tasks, measuring practical engineering skills across diverse software development scenarios.

What is the SkillsBench leaderboard?

The SkillsBench leaderboard ranks 8 AI models based on their performance on this benchmark. Currently, Qwen3.8 Max by Alibaba Cloud / Qwen Team leads with a score of 0.702. The average score across all models is 0.508.

What is the highest SkillsBench score?

The highest SkillsBench score is 0.702, achieved by Qwen3.8 Max from Alibaba Cloud / Qwen Team.

How many models are evaluated on SkillsBench?

8 models have been evaluated on the SkillsBench benchmark, with 0 verified results and 8 self-reported results.

What categories does SkillsBench cover?

SkillsBench is categorized under agents and code. The benchmark evaluates text models.

What is the best open-source model on SkillsBench?

Qwen3.8 Max by Alibaba Cloud / Qwen Team is the top-ranked open-source model on SkillsBench, with a score of 0.702 (rank #1).

Which model offers the best value on SkillsBench?

Among models scoring within 10% of the leader, Qwen3.8 Max from Alibaba Cloud / Qwen Team is the cheapest, at $1.65 per million input tokens with a score of 0.702.

How recent are the SkillsBench leaderboard results?

The SkillsBench leaderboard was last updated in August 2026 and currently includes 8 evaluated models.