GraphWalks
Progress Over Time
Interactive timeline showing model performance evolution on GraphWalks
GraphWalks Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Microsoft | 1.0T | — | — | ||
| 2 | Xiaomi | 311B | 1.0M | $0.17 / $0.34 | ||
| 3 | Xiaomi | 1.0T | 1.0M | $0.43 / $0.87 |
What is GraphWalks?
GraphWalks is a synthetic multi-hop long-context reasoning benchmark in which a model is given an edge-list representation of a graph and must traverse it to find neighboring nodes (via breadth-first search) or parent nodes for a given start node. Performance is reported as F1 of the model-predicted answer set versus the ground truth.
GraphWalks is a text benchmark evaluating models on long context and reasoning tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.
Compare leaders on the best AI for long context and best AI for reasoning leaderboards.
Current leaders
MAI-Thinking-1 from Microsoft currently leads the GraphWalks leaderboard with a score of 0.900 across 3 evaluated AI models.
FAQ
Common questions about the GraphWalks benchmark and leaderboard.