The AI arena is free today

Open Superagent

Claw-Eval

Progress Over Time

Interactive timeline showing model performance evolution on Claw-Eval

State-of-the-art frontier
Open
Proprietary

Claw-Eval Leaderboard

14 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
2
Zhipu AI
Zhipu AI
3
MiniMax
MiniMax
428B1.0M$0.30 / $1.20
4
Tencent
Tencent
295B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
61.0T1.0M$0.43 / $0.87
7
Xiaomi
Xiaomi
311B1.0M$0.17 / $0.34
8
Liquid AI
Liquid AI
3B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
101.0T
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B262K$0.60 / $3.60
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
13
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is Claw-Eval?

Claw-Eval tests real-world agentic task completion across complex multi-step scenarios, evaluating a model's ability to use tools, navigate environments, and complete end-to-end tasks autonomously.

Claw-Eval is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 14 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.8.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Kimi K2.6 from Moonshot AI currently leads the Claw-Eval leaderboard with a score of 0.809 across 14 evaluated AI models.

1Kimi K2.6Moonshot AI80.9%
2GLM-5V-TurboZhipu AI75.0%
3MiniMax M3MiniMax74.5%

FAQ

Common questions about the Claw-Eval benchmark and leaderboard.

What is the Claw-Eval benchmark?

Claw-Eval tests real-world agentic task completion across complex multi-step scenarios, evaluating a model's ability to use tools, navigate environments, and complete end-to-end tasks autonomously.

What is the Claw-Eval leaderboard?

The Claw-Eval leaderboard ranks 14 AI models based on their performance on this benchmark. Currently, Kimi K2.6 by Moonshot AI leads with a score of 0.809. The average score across all models is 0.645.

What is the highest Claw-Eval score?

The highest Claw-Eval score is 0.809, achieved by Kimi K2.6 from Moonshot AI.

How many models are evaluated on Claw-Eval?

14 models have been evaluated on the Claw-Eval benchmark, with 0 verified results and 14 self-reported results.

What categories does Claw-Eval cover?

Claw-Eval is categorized under agents and code. The benchmark evaluates text models.

Are there variants of Claw-Eval?

Yes. Claw-Eval has 1 related variant: Kimi Claw 24/7 Bench.

What is the best open-source model on Claw-Eval?

Kimi K2.6 by Moonshot AI is the top-ranked open-source model on Claw-Eval, with a score of 0.809 (rank #1).

Which model offers the best value on Claw-Eval?

Among models scoring within 10% of the leader, MiniMax M3 from MiniMax is the cheapest, at $0.30 per million input tokens with a score of 0.745.

How recent are the Claw-Eval leaderboard results?

The Claw-Eval leaderboard was last updated in August 2026 and currently includes 14 evaluated models.