The AI arena is free today

Open Superagent

CyberGym

Progress Over Time

Interactive timeline showing model performance evolution on CyberGym

State-of-the-art frontier
Open
Proprietary

CyberGym Leaderboard

16 models
ContextCostLicense
1763B1.0M$0.22 / $0.66
2
3
Zhipu AI
Zhipu AI
753B1.0M$1.20 / $4.00
41.6T1.0M$0.43 / $0.87
5
6
7
OpenAI
OpenAI
1.1M$5.00 / $30.00
81.0M$5.00 / $25.00
9770B
10304B1.0M$0.06 / $0.18
111.0M$5.00 / $25.00
121.0M$5.00 / $25.00
13
ByteDance
ByteDance
14
Zhipu AI
Zhipu AI
754B203K$1.05 / $3.50
15
16
Moonshot AI
Moonshot AI
1.0T
Notice missing or incorrect data?
About this benchmark

What is CyberGym?

CyberGym is a benchmark for evaluating AI agents on cybersecurity tasks, testing their ability to identify vulnerabilities, perform security analysis, and complete security-related challenges in a controlled environment.

CyberGym is a text benchmark evaluating models on safety, agents, and code tasks. LLM Stats tracks 16 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.

Compare leaders on the best AI for safety, best AI for agents and best AI for code leaderboards.

Current leaders

DeepSeek-V4.1-Flash from DeepSeek currently leads the CyberGym leaderboard with a score of 0.881 across 16 evaluated AI models.

1DeepSeek-V4.1-FlashDeepSeek88.1%
3GLM-5.3Zhipu AI84.5%

FAQ

Common questions about the CyberGym benchmark and leaderboard.

What is the CyberGym benchmark?

CyberGym is a benchmark for evaluating AI agents on cybersecurity tasks, testing their ability to identify vulnerabilities, perform security analysis, and complete security-related challenges in a controlled environment.

What is the CyberGym leaderboard?

The CyberGym leaderboard ranks 16 AI models based on their performance on this benchmark. Currently, DeepSeek-V4.1-Flash by DeepSeek leads with a score of 0.881. The average score across all models is 0.761.

What is the highest CyberGym score?

The highest CyberGym score is 0.881, achieved by DeepSeek-V4.1-Flash from DeepSeek.

How many models are evaluated on CyberGym?

16 models have been evaluated on the CyberGym benchmark, with 0 verified results and 16 self-reported results.

What categories does CyberGym cover?

CyberGym is categorized under safety, agents, and code. The benchmark evaluates text models.

What is the best open-source model on CyberGym?

DeepSeek-V4.1-Flash by DeepSeek is the top-ranked open-source model on CyberGym, with a score of 0.881 (rank #1).

Which model offers the best value on CyberGym?

Among models scoring within 10% of the leader, DeepSeek-V4.1-Flash from DeepSeek is the cheapest, at $0.22 per million input tokens with a score of 0.881.

How recent are the CyberGym leaderboard results?

The CyberGym leaderboard was last updated in September 2026 and currently includes 16 evaluated models.