The AI arena is free today

Open Superagent

CWE-Bench

Implementation

Progress Over Time

Interactive timeline showing model performance evolution on CWE-Bench

State-of-the-art frontier
Open
Proprietary

CWE-Bench Leaderboard

1 models
ContextCostLicense
1
Notice missing or incorrect data?
About this benchmark

What is CWE-Bench?

CWE-Bench evaluates autonomous agents on finding and patching real-world software vulnerabilities.

CWE-Bench is a text benchmark evaluating models on safety, agents, and code tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for safety, best AI for agents and best AI for code leaderboards.

Current leaders

Gemini 3.8 Flash Cyber from Google currently leads the CWE-Bench leaderboard with a score of 0.472 across 1 evaluated AI models.

FAQ

Common questions about the CWE-Bench benchmark and leaderboard.

What is the CWE-Bench benchmark?

CWE-Bench evaluates autonomous agents on finding and patching real-world software vulnerabilities.

What is the CWE-Bench leaderboard?

The CWE-Bench leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Gemini 3.8 Flash Cyber by Google leads with a score of 0.472. The average score across all models is 0.472.

What is the highest CWE-Bench score?

The highest CWE-Bench score is 0.472, achieved by Gemini 3.8 Flash Cyber from Google.

How many models are evaluated on CWE-Bench?

1 models have been evaluated on the CWE-Bench benchmark, with 0 verified results and 1 self-reported results.

Where can I find the CWE-Bench dataset?

The CWE-Bench dataset is available at https://cwe-bench.com/.

What categories does CWE-Bench cover?

CWE-Bench is categorized under safety, agents, and code. The benchmark evaluates text models.

How recent are the CWE-Bench leaderboard results?

The CWE-Bench leaderboard was last updated in September 2026 and currently includes 1 evaluated models.