CursorBench v3.2
Progress Over Time
Interactive timeline showing model performance evolution on CursorBench v3.2
State-of-the-art frontier
Open
Proprietary
CursorBench v3.2 Leaderboard
1 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | xAI | — | 500K | $2.00 / $6.00 |
Notice missing or incorrect data?
What is CursorBench v3.2?
CursorBench v3.2 evaluates coding agents on interactive software engineering tasks in the Cursor environment.
CursorBench v3.2 is a text benchmark evaluating models on agents and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for agents and best AI for code leaderboards.
Current leaders
Grok 4.6 from xAI currently leads the CursorBench v3.2 leaderboard with a score of 0.699 across 1 evaluated AI models.
FAQ
Common questions about the CursorBench v3.2 benchmark and leaderboard.