CL-bench
Progress Over Time
Interactive timeline showing model performance evolution on CL-bench
State-of-the-art frontier
Open
Proprietary
CL-bench Leaderboard
2 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Tencent | 295B | — | — | ||
| 2 | MiniMax | — | 1.0M | $0.30 / $1.20 |
Notice missing or incorrect data?
Sub-benchmarks
What is CL-bench?
CL-bench is an open-source benchmark with its own data and rubrics for evaluating models on coding and agentic tasks, scored using a setup fully aligned with the official procedure.
CL-bench is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.2.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Hy3 from Tencent currently leads the CL-bench leaderboard with a score of 0.238 across 2 evaluated AI models.
FAQ
Common questions about the CL-bench benchmark and leaderboard.