The AI arena is free today

Open Superagent

Graphwalks BFS >128k

Progress Over Time

Interactive timeline showing model performance evolution on Graphwalks BFS >128k

State-of-the-art frontier
Open
Proprietary

Graphwalks BFS >128k Leaderboard

11 models
ContextCostLicense
11.1M$5.00 / $30.00
21.1M$0.20 / $1.20
3
41.1M$2.00 / $12.00
51.0M$5.00 / $25.00
61.0M$5.00 / $25.00
7
OpenAI
OpenAI
1.1M$5.00 / $30.00
8
OpenAI
OpenAI
1.0M$2.50 / $15.00
9
OpenAI
OpenAI
1.0M$2.00 / $8.00
101.0M$0.40 / $1.60
111.0M$0.10 / $0.40
Notice missing or incorrect data?

Sub-benchmarks

About this benchmark

What is Graphwalks BFS >128k?

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.

Graphwalks BFS >128k is a text benchmark evaluating models on long context, reasoning, and spatial reasoning tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.9.

Compare leaders on the best AI for long context, best AI for reasoning and best AI for spatial reasoning leaderboards.

Current leaders

GPT-5.6 Sol from OpenAI currently leads the Graphwalks BFS >128k leaderboard with a score of 0.907 across 11 evaluated AI models.

1GPT-5.6 SolOpenAI90.7%
2GPT-5.6 LunaOpenAI81.3%
3Claude Mythos PreviewAnthropic80.0%

FAQ

Common questions about the Graphwalks BFS >128k benchmark and leaderboard.

What is the Graphwalks BFS >128k benchmark?

A graph reasoning benchmark that evaluates language models' ability to perform breadth-first search (BFS) operations on graphs with context length over 128k tokens, testing long-context reasoning capabilities.

What is the Graphwalks BFS >128k leaderboard?

The Graphwalks BFS >128k leaderboard ranks 11 AI models based on their performance on this benchmark. Currently, GPT-5.6 Sol by OpenAI leads with a score of 0.907. The average score across all models is 0.511.

What is the highest Graphwalks BFS >128k score?

The highest Graphwalks BFS >128k score is 0.907, achieved by GPT-5.6 Sol from OpenAI.

How many models are evaluated on Graphwalks BFS >128k?

11 models have been evaluated on the Graphwalks BFS >128k benchmark, with 0 verified results and 11 self-reported results.

What categories does Graphwalks BFS >128k cover?

Graphwalks BFS >128k is categorized under long context, reasoning, and spatial reasoning. The benchmark evaluates text models.

Are there variants of Graphwalks BFS >128k?

Yes. Graphwalks BFS >128k has 1 related variant: Graphwalks BFS 1M.

Which model offers the best value on Graphwalks BFS >128k?

Among models scoring within 10% of the leader, GPT-5.6 Sol from OpenAI is the cheapest, at $5.00 per million input tokens with a score of 0.907.

How recent are the Graphwalks BFS >128k leaderboard results?

The Graphwalks BFS >128k leaderboard was last updated in September 2026 and currently includes 11 evaluated models.