APEX-Agents
Progress Over Time
Interactive timeline showing model performance evolution on APEX-Agents
APEX-Agents Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | ByteDance | — | — | — | ||
| 3 | Google | — | 1.0M | $2.50 / $15.00 | ||
| 4 | ByteDance | — | — | — | ||
| 5 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 6 | MiniMax | — | 1.0M | $0.30 / $1.20 | ||
| 7 | Tencent | 295B | — | — |
What is APEX-Agents?
APEX-Agents is a benchmark evaluating AI agents on long horizon professional tasks that require sustained reasoning, planning, and execution across complex multi-step workflows.
APEX-Agents is a text benchmark evaluating models on reasoning and agents tasks. LLM Stats tracks 7 models on this benchmark, scored on a 0–1 scale. The current average is 0.3, with the leader at 0.4.
Compare leaders on the best AI for reasoning and best AI for agents leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the APEX-Agents leaderboard with a score of 0.376 across 7 evaluated AI models.
FAQ
Common questions about the APEX-Agents benchmark and leaderboard.