MCP-Mark
Progress Over Time
Interactive timeline showing model performance evolution on MCP-Mark
MCP-Mark Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 3 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 4 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 5 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 6 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 7 | DeepSeek | 685B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 35B | — | — |
What is MCP-Mark?
MCP-Mark evaluates LLMs on their ability to use Model Context Protocol (MCP) tools effectively, testing tool discovery, selection, invocation, and result interpretation across diverse MCP server scenarios.
MCP-Mark is a text benchmark evaluating models on agents and tool calling tasks. LLM Stats tracks 8 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.8.
Compare leaders on the best AI for agents and best AI for tool calling leaderboards.
Current leaders
Kimi K2.7 Code from Moonshot AI currently leads the MCP-Mark leaderboard with a score of 0.811 across 8 evaluated AI models.
FAQ
Common questions about the MCP-Mark benchmark and leaderboard.