The AI arena is free today

Open Superagent

MCP Atlas

Progress Over Time

Interactive timeline showing model performance evolution on MCP Atlas

State-of-the-art frontier
Open
Proprietary

MCP Atlas Leaderboard

36 models
ContextCostLicense
11.0M$1.25 / $4.25
2
Moonshot AI
Moonshot AI
2.8T1.0M$2.85 / $14.25
3
ByteDance
ByteDance
4770B
51.0M$1.50 / $9.00
61.0M$5.00 / $25.00
7
8
Thinking Machines Lab
Thinking Machines Lab
276B524K$0.30 / $1.20
9
Tencent
Tencent
295B262K$0.14 / $0.58
101.0M$5.00 / $25.00
11
Zhipu AI
Zhipu AI
753B1.0M$0.75 / $2.40
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
13
Moonshot AI
Moonshot AI
1.0T262K$0.68 / $3.40
13
Thinking Machines Lab
Thinking Machines Lab
975B524K$0.95 / $4.05
1530B131K$0.30 / $1.20
16
OpenAI
OpenAI
1.1M$5.00 / $30.00
17
MiniMax
MiniMax
428B1.0M$0.28 / $1.10
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
191.6T1.0M$1.30 / $2.60
20
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
21
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
221.0M$2.00 / $12.00
23284B1.0M$0.09 / $0.18
24
Zhipu AI
Zhipu AI
744B200K$1.00 / $3.20
25284B1.0M$0.09 / $0.18
26
OpenAI
OpenAI
1.0M$2.50 / $15.00
27
InclusionAI
InclusionAI
124B131K$0.06 / $0.18
28
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B262K$0.10 / $0.95
291.0M$5.00 / $25.00
30
311.0M$3.00 / $15.00
32
OpenAI
OpenAI
400K$1.75 / $14.00
33400K$0.75 / $4.50
341.0M$0.50 / $3.00
35400K$0.20 / $1.25
361.0M$0.30 / $2.50
Notice missing or incorrect data?
About this benchmark

What is MCP Atlas?

MCP Atlas is a benchmark for evaluating AI models on scaled tool use capabilities, measuring how well models can coordinate and utilize multiple tools across complex multi-step tasks.

MCP Atlas is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 36 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.9.

Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.

Current leaders

Muse Spark 1.1 from Meta currently leads the MCP Atlas leaderboard with a score of 0.881 across 36 evaluated AI models.

1Muse Spark 1.1Meta88.1%
2Kimi K3Moonshot AI84.2%
3Seed 2.1 ProByteDance83.8%
OSSHy4 preview#4 open-weight83.7%

FAQ

Common questions about the MCP Atlas benchmark and leaderboard.

What is the MCP Atlas benchmark?

MCP Atlas is a benchmark for evaluating AI models on scaled tool use capabilities, measuring how well models can coordinate and utilize multiple tools across complex multi-step tasks.

What is the MCP Atlas leaderboard?

The MCP Atlas leaderboard ranks 36 AI models based on their performance on this benchmark. Currently, Muse Spark 1.1 by Meta leads with a score of 0.881. The average score across all models is 0.710.

What is the highest MCP Atlas score?

The highest MCP Atlas score is 0.881, achieved by Muse Spark 1.1 from Meta.

How many models are evaluated on MCP Atlas?

36 models have been evaluated on the MCP Atlas benchmark, with 0 verified results and 36 self-reported results.

What categories does MCP Atlas cover?

MCP Atlas is categorized under reasoning, agents, code, and tool calling. The benchmark evaluates text models.

What is the best open-source model on MCP Atlas?

Hy4 preview by Tencent is the top-ranked open-source model on MCP Atlas, with a score of 0.837 (rank #4).

Which model offers the best value on MCP Atlas?

Among models scoring within 10% of the leader, Inkling-Small from Thinking Machines Lab is the cheapest, at $0.30 per million input tokens with a score of 0.796.

How recent are the MCP Atlas leaderboard results?

The MCP Atlas leaderboard was last updated in September 2026 and currently includes 36 evaluated models.