Gray Swan IPI Benchmark
Progress Over Time
Interactive timeline showing model performance evolution on Gray Swan IPI Benchmark
Gray Swan IPI Benchmark Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | — | — |
What is Gray Swan IPI Benchmark?
The Gray Swan IPI Benchmark evaluates indirect prompt-injection robustness by measuring attack success within a fixed number of attempts; lower is better.
Gray Swan IPI Benchmark is a text benchmark evaluating models on safety, agents, and tool calling tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.1, with the leader at 0.1.
Compare leaders on the best AI for safety, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Gemini 3.8 Flash Cyber from Google currently leads the Gray Swan IPI Benchmark leaderboard with a score of 0.060 across 1 evaluated AI models.
FAQ
Common questions about the Gray Swan IPI Benchmark benchmark and leaderboard.