MCP Atlas
Progress Over Time
Interactive timeline showing model performance evolution on MCP Atlas
MCP Atlas Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Meta | — | 1.0M | $1.25 / $4.25 | ||
| 2 | Moonshot AI | 2.8T | 1.0M | $2.85 / $14.25 | ||
| 3 | ByteDance | — | — | — | ||
| 4 | Tencent | 770B | — | — | ||
| 5 | Google | — | 1.0M | $1.50 / $9.00 | ||
| 6 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 7 | ByteDance | — | — | — | ||
| 8 | Thinking Machines Lab | 276B | 524K | $0.30 / $1.20 | ||
| 9 | Tencent | 295B | 262K | $0.14 / $0.58 | ||
| 10 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 11 | Zhipu AI | 753B | 1.0M | $0.75 / $2.40 | ||
| 12 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 13 | Moonshot AI | 1.0T | 262K | $0.68 / $3.40 | ||
| 13 | Thinking Machines Lab | 975B | 524K | $0.95 / $4.05 | ||
| 15 | Meta | 30B | 131K | $0.30 / $1.20 | ||
| 16 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 17 | MiniMax | 428B | 1.0M | $0.28 / $1.10 | ||
| 18 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 19 | DeepSeek | 1.6T | 1.0M | $1.30 / $2.60 | ||
| 20 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 21 | Zhipu AI | 754B | 203K | $1.05 / $3.50 | ||
| 22 | Google | — | 1.0M | $2.00 / $12.00 | ||
| 23 | DeepSeek | 284B | 1.0M | $0.09 / $0.18 | ||
| 24 | Zhipu AI | 744B | 200K | $1.00 / $3.20 | ||
| 25 | DeepSeek | 284B | 1.0M | $0.09 / $0.18 | ||
| 26 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 27 | InclusionAI | 124B | 131K | $0.06 / $0.18 | ||
| 28 | Alibaba Cloud / Qwen Team | 35B | 262K | $0.10 / $0.95 | ||
| 29 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 30 | Anthropic | — | — | — | ||
| 31 | Anthropic | — | 1.0M | $3.00 / $15.00 | ||
| 32 | OpenAI | — | 400K | $1.75 / $14.00 | ||
| 33 | OpenAI | — | 400K | $0.75 / $4.50 | ||
| 34 | Google | — | 1.0M | $0.50 / $3.00 | ||
| 35 | OpenAI | — | 400K | $0.20 / $1.25 | ||
| 36 | Amazon | — | 1.0M | $0.30 / $2.50 |
What is MCP Atlas?
MCP Atlas is a benchmark for evaluating AI models on scaled tool use capabilities, measuring how well models can coordinate and utilize multiple tools across complex multi-step tasks.
MCP Atlas is a text benchmark evaluating models on reasoning, agents, code, and tool calling tasks. LLM Stats tracks 36 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.9.
Compare leaders on the best AI for reasoning, best AI for agents, best AI for code and best AI for tool calling leaderboards.
Current leaders
Muse Spark 1.1 from Meta currently leads the MCP Atlas leaderboard with a score of 0.881 across 36 evaluated AI models.
FAQ
Common questions about the MCP Atlas benchmark and leaderboard.