The AI arena is free today

Open Superagent

Gray Swan IPI Benchmark

Progress Over Time

Interactive timeline showing model performance evolution on Gray Swan IPI Benchmark

State-of-the-art frontier
Open
Proprietary

Gray Swan IPI Benchmark Leaderboard

1 models
ContextCostLicense
1
Notice missing or incorrect data?
About this benchmark

What is Gray Swan IPI Benchmark?

The Gray Swan IPI Benchmark evaluates indirect prompt-injection robustness by measuring attack success within a fixed number of attempts; lower is better.

Gray Swan IPI Benchmark is a text benchmark evaluating models on safety, agents, and tool calling tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.1, with the leader at 0.1.

Compare leaders on the best AI for safety, best AI for agents and best AI for tool calling leaderboards.

Current leaders

Gemini 3.8 Flash Cyber from Google currently leads the Gray Swan IPI Benchmark leaderboard with a score of 0.060 across 1 evaluated AI models.

FAQ

Common questions about the Gray Swan IPI Benchmark benchmark and leaderboard.

What is the Gray Swan IPI Benchmark benchmark?

The Gray Swan IPI Benchmark evaluates indirect prompt-injection robustness by measuring attack success within a fixed number of attempts; lower is better.

What is the Gray Swan IPI Benchmark leaderboard?

The Gray Swan IPI Benchmark leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Gemini 3.8 Flash Cyber by Google leads with a score of 0.060. The average score across all models is 0.060.

What is the highest Gray Swan IPI Benchmark score?

The highest Gray Swan IPI Benchmark score is 0.060, achieved by Gemini 3.8 Flash Cyber from Google.

How many models are evaluated on Gray Swan IPI Benchmark?

1 models have been evaluated on the Gray Swan IPI Benchmark benchmark, with 0 verified results and 1 self-reported results.

What categories does Gray Swan IPI Benchmark cover?

Gray Swan IPI Benchmark is categorized under safety, agents, and tool calling. The benchmark evaluates text models.

How recent are the Gray Swan IPI Benchmark leaderboard results?

The Gray Swan IPI Benchmark leaderboard was last updated in September 2026 and currently includes 1 evaluated models.