MASK
Progress Over Time
Interactive timeline showing model performance evolution on MASK
MASK Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | — | — | — |
What is MASK?
MASK is a collection of 1000 questions measuring whether models faithfully report their beliefs when pressured to lie. It operationalizes deception as the rate at which the model lies, i.e., knowingly making false statements intended to be received as true. Lower dishonesty rates indicate better honesty.
MASK is a text benchmark evaluating models on reasoning and safety tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.
Compare leaders on the best AI for reasoning and best AI for safety leaderboards.
Current leaders
Grok-4.1 Thinking from xAI currently leads the MASK leaderboard with a score of 0.510 across 1 evaluated AI models.
FAQ
Common questions about the MASK benchmark and leaderboard.