The AI arena is free today

Open Superagent

APEX-Agents

Progress Over Time

Interactive timeline showing model performance evolution on APEX-Agents

State-of-the-art frontier
Open
Proprietary

APEX-Agents Leaderboard

10 models
ContextCostLicense
1500K$2.00 / $6.00
2
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
3770B
4
ByteDance
ByteDance
51.0M$2.00 / $12.00
6
7124B262K$0.06 / $0.18
8
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
9
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
10
Tencent
Tencent
295B262K$0.14 / $0.58
Notice missing or incorrect data?
About this benchmark

What is APEX-Agents?

APEX-Agents is a benchmark evaluating AI agents on long horizon professional tasks that require sustained reasoning, planning, and execution across complex multi-step workflows.

APEX-Agents is a text benchmark evaluating models on reasoning and agents tasks. LLM Stats tracks 10 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.6.

Compare leaders on the best AI for reasoning and best AI for agents leaderboards.

Current leaders

Grok 4.6 from xAI currently leads the APEX-Agents leaderboard with a score of 0.575 across 10 evaluated AI models.

1Grok 4.6xAI57.5%
2Kimi K3Moonshot AI37.6%
3Hy4 previewTencent37.1%

FAQ

Common questions about the APEX-Agents benchmark and leaderboard.

What is the APEX-Agents benchmark?

APEX-Agents is a benchmark evaluating AI agents on long horizon professional tasks that require sustained reasoning, planning, and execution across complex multi-step workflows.

What is the APEX-Agents leaderboard?

The APEX-Agents leaderboard ranks 10 AI models based on their performance on this benchmark. Currently, Grok 4.6 by xAI leads with a score of 0.575. The average score across all models is 0.339.

What is the highest APEX-Agents score?

The highest APEX-Agents score is 0.575, achieved by Grok 4.6 from xAI.

How many models are evaluated on APEX-Agents?

10 models have been evaluated on the APEX-Agents benchmark, with 0 verified results and 10 self-reported results.

What categories does APEX-Agents cover?

APEX-Agents is categorized under reasoning and agents. The benchmark evaluates text models.

What is the best open-source model on APEX-Agents?

Hy4 preview by Tencent is the top-ranked open-source model on APEX-Agents, with a score of 0.371 (rank #3).

Which model offers the best value on APEX-Agents?

Among models scoring within 10% of the leader, Grok 4.6 from xAI is the cheapest, at $2.00 per million input tokens with a score of 0.575.

How recent are the APEX-Agents leaderboard results?

The APEX-Agents leaderboard was last updated in September 2026 and currently includes 10 evaluated models.