Siren AgentDojo Utility
Progress Over Time
Interactive timeline showing model performance evolution on Siren AgentDojo Utility
Siren AgentDojo Utility Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Meta | 30B | — | — |
What is Siren AgentDojo Utility?
Siren AgentDojo evaluates tool-using agents under adversarial prompt-injection attacks. This metric reports utility on the assigned tasks.
Siren AgentDojo Utility is a text benchmark evaluating models on safety, agents, and tool calling tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.
Compare leaders on the best AI for safety, best AI for agents and best AI for tool calling leaderboards.
Current leaders
Muse Glimmer-30B from Meta currently leads the Siren AgentDojo Utility leaderboard with a score of 0.942 across 1 evaluated AI models.
FAQ
Common questions about the Siren AgentDojo Utility benchmark and leaderboard.