The AI arena is free today

Open Superagent

PerceptionTest

Paper

Progress Over Time

Interactive timeline showing model performance evolution on PerceptionTest

State-of-the-art frontier
Open
Proprietary

PerceptionTest Leaderboard

2 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
72B
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
Notice missing or incorrect data?
About this benchmark

What is PerceptionTest?

A novel multimodal video benchmark designed to evaluate perception and reasoning skills of pre-trained models across video, audio, and text modalities. Contains 11.6k real-world videos (average 23 seconds) filmed by participants worldwide, densely annotated with six types of labels. Focuses on skills (Memory, Abstraction, Physics, Semantics) and reasoning types (descriptive, explanatory, predictive, counterfactual). Shows significant performance gap between human baseline (91.4%) and state-of-the-art video QA models (46.2%).

PerceptionTest is a multimodal benchmark evaluating models on multimodal, physics, reasoning, spatial reasoning, video, and vision tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.7.

Compare leaders on the best AI for multimodal, best AI for physics, best AI for reasoning, best AI for spatial reasoning, best AI for video and best AI for vision leaderboards.

Current leaders

Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team currently leads the PerceptionTest leaderboard with a score of 0.732 across 2 evaluated AI models.

1Qwen2.5 VL 72B InstructAlibaba Cloud / Qwen Team73.2%
2Qwen2.5 VL 7B InstructAlibaba Cloud / Qwen Team70.5%

Source paper

Title
Perception Test: A Diagnostic Benchmark for Multimodal Video Models
Authors
Viorica Pătrăucean, Lucas Smaira, Ankush Gupta, Adrià Recasens Continente, and 20 others
Published
Abstract

We propose a novel multimodal video benchmark - the Perception Test - to evaluate the perception and reasoning skills of pre-trained multimodal models (e.g. Flamingo, SeViLA, or GPT-4). Compared to existing benchmarks that focus on computational tasks (e.g. classification, detection or tracking), the Perception Test focuses on skills (Memory, Abstraction, Physics, Semantics) and types of reasoning (descriptive, explanatory, predictive, counterfactual) across video, audio, and text modalities, to provide a comprehensive and efficient evaluation tool. The benchmark probes pre-trained models for their transfer capabilities, in a zero-shot / few-shot or limited finetuning regime. For these purposes, the Perception Test introduces 11.6k real-world videos, 23s average length, designed to show perceptually interesting situations, filmed by around 100 participants worldwide. The videos are densely annotated with six types of labels (multiple-choice and grounded video question-answers, object and point tracks, temporal action and sound segments), enabling both language and non-language evaluations. The fine-tuning and validation splits of the benchmark are publicly available (CC-BY license), in addition to a challenge server with a held-out test split. Human baseline results compared to state-of-the-art video QA models show a substantial gap in performance (91.4% vs 46.2%), suggesting that there is significant room for improvement in multimodal video understanding. Dataset, baseline code, and challenge server are available at https://github.com/deepmind/perception_test

FAQ

Common questions about the PerceptionTest benchmark and leaderboard.

What is the PerceptionTest benchmark?

A novel multimodal video benchmark designed to evaluate perception and reasoning skills of pre-trained models across video, audio, and text modalities. Contains 11.6k real-world videos (average 23 seconds) filmed by participants worldwide, densely annotated with six types of labels. Focuses on skills (Memory, Abstraction, Physics, Semantics) and reasoning types (descriptive, explanatory, predictive, counterfactual). Shows significant performance gap between human baseline (91.4%) and state-of-the-art video QA models (46.2%).

What is the PerceptionTest leaderboard?

The PerceptionTest leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Qwen2.5 VL 72B Instruct by Alibaba Cloud / Qwen Team leads with a score of 0.732. The average score across all models is 0.718.

What is the highest PerceptionTest score?

The highest PerceptionTest score is 0.732, achieved by Qwen2.5 VL 72B Instruct from Alibaba Cloud / Qwen Team.

How many models are evaluated on PerceptionTest?

2 models have been evaluated on the PerceptionTest benchmark, with 0 verified results and 2 self-reported results.

Where can I find the PerceptionTest paper?

The PerceptionTest paper is available at https://arxiv.org/abs/2305.13786. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does PerceptionTest cover?

PerceptionTest is categorized under multimodal, physics, reasoning, spatial reasoning, video, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on PerceptionTest?

Qwen2.5 VL 7B Instruct by Alibaba Cloud / Qwen Team is the top-ranked open-source model on PerceptionTest, with a score of 0.705 (rank #2).

How recent are the PerceptionTest leaderboard results?

The PerceptionTest leaderboard was last updated in August 2026 and currently includes 2 evaluated models.