The AI arena is free today

Open Superagent

MASK

Paper

Progress Over Time

Interactive timeline showing model performance evolution on MASK

State-of-the-art frontier
Open
Proprietary

MASK Leaderboard

1 models
ContextCostLicense
1
Notice missing or incorrect data?
About this benchmark

What is MASK?

MASK is a collection of 1000 questions measuring whether models faithfully report their beliefs when pressured to lie. It operationalizes deception as the rate at which the model lies, i.e., knowingly making false statements intended to be received as true. Lower dishonesty rates indicate better honesty.

MASK is a text benchmark evaluating models on reasoning and safety tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for reasoning and best AI for safety leaderboards.

Current leaders

Grok-4.1 Thinking from xAI currently leads the MASK leaderboard with a score of 0.510 across 1 evaluated AI models.

FAQ

Common questions about the MASK benchmark and leaderboard.

What is the MASK benchmark?

MASK is a collection of 1000 questions measuring whether models faithfully report their beliefs when pressured to lie. It operationalizes deception as the rate at which the model lies, i.e., knowingly making false statements intended to be received as true. Lower dishonesty rates indicate better honesty.

What is the MASK leaderboard?

The MASK leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Grok-4.1 Thinking by xAI leads with a score of 0.510. The average score across all models is 0.510.

What is the highest MASK score?

The highest MASK score is 0.510, achieved by Grok-4.1 Thinking from xAI.

How many models are evaluated on MASK?

1 models have been evaluated on the MASK benchmark, with 0 verified results and 1 self-reported results.

Where can I find the MASK paper?

The MASK paper is available at https://arxiv.org/abs/2503.03750. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does MASK cover?

MASK is categorized under reasoning and safety. The benchmark evaluates text models.

How recent are the MASK leaderboard results?

The MASK leaderboard was last updated in August 2026 and currently includes 1 evaluated models.