The AI arena is free today

Open Superagent

CursorBench v3.2

Progress Over Time

Interactive timeline showing model performance evolution on CursorBench v3.2

State-of-the-art frontier
Open
Proprietary

CursorBench v3.2 Leaderboard

1 models
ContextCostLicense
1500K$2.00 / $6.00
Notice missing or incorrect data?
About this benchmark

What is CursorBench v3.2?

CursorBench v3.2 evaluates coding agents on interactive software engineering tasks in the Cursor environment.

CursorBench v3.2 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for agents and best AI for code leaderboards.

Current leaders

Grok 4.6 from xAI currently leads the CursorBench v3.2 leaderboard with a score of 0.699 across 1 evaluated AI models.

1Grok 4.6xAI69.9%

FAQ

Common questions about the CursorBench v3.2 benchmark and leaderboard.

What is the CursorBench v3.2 benchmark?

CursorBench v3.2 evaluates coding agents on interactive software engineering tasks in the Cursor environment.

What is the CursorBench v3.2 leaderboard?

The CursorBench v3.2 leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Grok 4.6 by xAI leads with a score of 0.699. The average score across all models is 0.699.

What is the highest CursorBench v3.2 score?

The highest CursorBench v3.2 score is 0.699, achieved by Grok 4.6 from xAI.

How many models are evaluated on CursorBench v3.2?

1 models have been evaluated on the CursorBench v3.2 benchmark, with 0 verified results and 1 self-reported results.

What categories does CursorBench v3.2 cover?

CursorBench v3.2 is categorized under agents and code. The benchmark evaluates text models.

Which model offers the best value on CursorBench v3.2?

Among models scoring within 10% of the leader, Grok 4.6 from xAI is the cheapest, at $2.00 per million input tokens with a score of 0.699.

How recent are the CursorBench v3.2 leaderboard results?

The CursorBench v3.2 leaderboard was last updated in August 2026 and currently includes 1 evaluated models.