Graphwalks BFS >128k
Progress Over Time
Interactive timeline showing model performance evolution on Graphwalks BFS >128k
Graphwalks BFS >128k Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 2 | OpenAI | — | 1.1M | $0.20 / $1.20 | ||
| 3 | Anthropic | — | — | — | ||
| 4 | OpenAI | — | 1.1M | $2.00 / $12.00 | ||
| 5 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 6 | Anthropic | — | 1.0M | $5.00 / $25.00 | ||
| 7 | OpenAI | — | 1.1M | $5.00 / $30.00 | ||
| 8 | OpenAI | — | 1.0M | $2.50 / $15.00 | ||
| 9 | OpenAI | — | 1.0M | $2.00 / $8.00 | ||
| 10 | OpenAI | — | 1.0M | $0.40 / $1.60 | ||
| 11 | OpenAI | — | 1.0M | $0.10 / $0.40 |
Sub-benchmarks
What is Graphwalks BFS >128k?
A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.
Graphwalks BFS >128k is a text benchmark evaluating models on long context, reasoning, and spatial reasoning tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.9.
Compare leaders on the best AI for long context, best AI for reasoning and best AI for spatial reasoning leaderboards.
Current leaders
GPT-5.6 Sol from OpenAI currently leads the Graphwalks BFS >128k leaderboard with a score of 0.907 across 11 evaluated AI models.
FAQ
Common questions about the Graphwalks BFS >128k benchmark and leaderboard.