The AI arena is free today

Open Superagent

WildClawBench

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on WildClawBench

State-of-the-art frontier
Open
Proprietary

WildClawBench Leaderboard

6 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
———
2———
3
ByteDance
ByteDance
———
4
Tencent
Tencent
295B262K$0.14 / $0.58
530B131K$0.30 / $1.20
61.0T1.0M$0.43 / $0.87
Notice missing or incorrect data?
About this benchmark

What is WildClawBench?

WildClawBench is an agentic coding benchmark from InternLM/Claw-Eval that reports overall model performance on real-world tool-using development tasks.

WildClawBench is a text benchmark evaluating models on agents and coding tasks. LLM Stats tracks 6 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for coding leaderboards.

Current leaders

Qwen3.8 Omni Flash from Alibaba Cloud / Qwen Team currently leads the WildClawBench leaderboard with a score of 0.710 across 6 evaluated AI models.

1Qwen3.8 Omni FlashAlibaba Cloud / Qwen Team71.0%
2Seed 2.1 TurboByteDance62.8%
3Seed 2.1 ProByteDance61.7%
OSSHy3#4 open-weight53.6%

FAQ

Common questions about the WildClawBench benchmark and leaderboard.

What is the WildClawBench benchmark?

WildClawBench is an agentic coding benchmark from InternLM/Claw-Eval that reports overall model performance on real-world tool-using development tasks.

What is the WildClawBench leaderboard?

The WildClawBench leaderboard ranks 6 AI models based on their performance on this benchmark. Currently, Qwen3.8 Omni Flash by Alibaba Cloud / Qwen Team leads with a score of 0.710. The average score across all models is 0.566.

What is the highest WildClawBench score?

The highest WildClawBench score is 0.710, achieved by Qwen3.8 Omni Flash from Alibaba Cloud / Qwen Team.

How many models are evaluated on WildClawBench?

6 models have been evaluated on the WildClawBench benchmark, with 0 verified results and 6 self-reported results.

Where can I find the WildClawBench dataset?

The WildClawBench dataset is available on HuggingFace at https://huggingface.co/datasets/internlm/WildClawBench.

What categories does WildClawBench cover?

WildClawBench is categorized under agents and coding. The benchmark evaluates text models.

What is the best open-source model on WildClawBench?

Hy3 by Tencent is the top-ranked open-source model on WildClawBench, with a score of 0.536 (rank #4).

How recent are the WildClawBench leaderboard results?

The WildClawBench leaderboard was last updated in October 2026 and currently includes 6 evaluated models.