ACEBench

PaperImplementation

Progress Over Time

Interactive timeline showing model performance evolution on ACEBench

State-of-the-art frontier
Open
Proprietary

ACEBench Leaderboard

2 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
1.0T
11.0T
Notice missing or incorrect data?
About this benchmark

What is ACEBench?

ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchmark covers 4,538 APIs across 8 major domains and 68 sub-domains including technology, finance, entertainment, society, health, culture, and environment, supporting both English and Chinese languages.

ACEBench is a text benchmark evaluating models on reasoning, finance, general, healthcare, and tool calling tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.8.

Compare leaders on the best AI for reasoning, best AI for finance, best AI for general, best AI for healthcare and best AI for tool calling leaderboards.

Current leaders

Kimi K2 Instruct from Moonshot AI currently leads the ACEBench leaderboard with a score of 0.765 across 2 evaluated AI models.

1Kimi K2 InstructMoonshot AI76.5%
1Kimi K2-Instruct-0905Moonshot AI76.5%

Source paper

Title
ACEBench: Who Wins the Match Point in Tool Usage?
Authors
Chen Chen, Xinlong Hao, Weiwen Liu, Xu Huang, and 12 others
Published
Abstract

Large Language Models (LLMs) have demonstrated significant potential in decision-making and reasoning, particularly when integrated with various tools to effectively solve complex problems. However, existing benchmarks for evaluating LLMs' tool usage face several limitations: (1) limited evaluation scenarios, often lacking assessments in real multi-turn dialogue contexts; (2) narrow evaluation dimensions, with insufficient detailed assessments of how LLMs use tools; and (3) reliance on LLMs or real API executions for evaluation, which introduces significant overhead. To address these challenges, we introduce ACEBench, a comprehensive benchmark for assessing tool usage in LLMs. ACEBench categorizes data into three primary types based on evaluation methodology: Normal, Special, and Agent. "Normal" evaluates tool usage in basic scenarios; "Special" evaluates tool usage in situations with ambiguous or incomplete instructions; "Agent" evaluates tool usage through multi-agent interactions to simulate real-world, multi-turn dialogues. We conducted extensive experiments using ACEBench, analyzing various LLMs in-depth and providing a more granular examination of error causes across different data types.

FAQ

Common questions about the ACEBench benchmark and leaderboard.

What is the ACEBench benchmark?

ACEBench is a comprehensive benchmark for evaluating Large Language Models' tool usage capabilities across three primary evaluation types: Normal (basic tool usage scenarios), Special (tool usage with ambiguous or incomplete instructions), and Agent (multi-agent interactions simulating real-world dialogues). The benchmark covers 4,538 APIs across 8 major domains and 68 sub-domains including technology, finance, entertainment, society, health, culture, and environment, supporting both English and Chinese languages.

What is the ACEBench leaderboard?

The ACEBench leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Kimi K2 Instruct by Moonshot AI leads with a score of 0.765. The average score across all models is 0.765.

What is the highest ACEBench score?

The highest ACEBench score is 0.765, achieved by Kimi K2 Instruct from Moonshot AI.

How many models are evaluated on ACEBench?

2 models have been evaluated on the ACEBench benchmark, with 0 verified results and 2 self-reported results.

Where can I find the ACEBench paper?

The ACEBench paper is available at https://arxiv.org/abs/2501.12851. The paper details the methodology, dataset construction, and evaluation criteria.

Where can I find the ACEBench dataset?

The ACEBench dataset is available at https://github.com/ACEBench/ACEBench.

What categories does ACEBench cover?

ACEBench is categorized under reasoning, finance, general, healthcare, and tool calling. The benchmark evaluates text models.

What is the best open-source model on ACEBench?

Kimi K2 Instruct by Moonshot AI is the top-ranked open-source model on ACEBench, with a score of 0.765 (rank #1).

How recent are the ACEBench leaderboard results?

The ACEBench leaderboard was last updated in July 2026 and currently includes 2 evaluated models.