The AI arena is free today

Open Superagent

LVBench

Paper

Progress Over Time

Interactive timeline showing model performance evolution on LVBench

State-of-the-art frontier
Open
Proprietary

LVBench Leaderboard

26 models
ContextCostLicense
11.0M$0.75 / $3.75
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2.4T1.0M$1.65 / $4.95
3
ByteDance
ByteDance
4
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
125B
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
7
Moonshot AI
Moonshot AI
1.0T
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
12
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
13
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
14
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
236B
15
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
33B
16
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
17
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
31B
18
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
19
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $0.60
20
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
21
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B262K$0.10 / $1.00
22
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
34B
23
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
72B
24
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
8B
25
Amazon
Amazon
26
Amazon
Amazon
Notice missing or incorrect data?
About this benchmark

What is LVBench?

LVBench is an extreme long video understanding benchmark designed to evaluate multimodal models on videos up to two hours in duration. It contains 6 major categories and 21 subcategories, with videos averaging five times longer than existing datasets. The benchmark addresses applications requiring comprehension of extremely long videos.

LVBench is a multimodal benchmark evaluating models on long context, multimodal, and vision tasks. LLM Stats tracks 26 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.9.

Compare leaders on the best AI for long context, best AI for multimodal and best AI for vision leaderboards.

Current leaders

Gemini 3.7 Flash from Google currently leads the LVBench leaderboard with a score of 0.854 across 26 evaluated AI models.

1Gemini 3.7 FlashGoogle85.4%
2Qwen3.8 MaxAlibaba Cloud / Qwen Team81.8%
3Seed 2.1 ProByteDance78.0%
OSSKimi K2.5#7 open-weight75.9%

Source paper

Title
LVBench: An Extreme Long Video Understanding Benchmark
Authors
Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, and 8 others
Published
Abstract

Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io.

FAQ

Common questions about the LVBench benchmark and leaderboard.

What is the LVBench benchmark?

LVBench is an extreme long video understanding benchmark designed to evaluate multimodal models on videos up to two hours in duration. It contains 6 major categories and 21 subcategories, with videos averaging five times longer than existing datasets. The benchmark addresses applications requiring comprehension of extremely long videos.

What is the LVBench leaderboard?

The LVBench leaderboard ranks 26 AI models based on their performance on this benchmark. Currently, Gemini 3.7 Flash by Google leads with a score of 0.854. The average score across all models is 0.642.

What is the highest LVBench score?

The highest LVBench score is 0.854, achieved by Gemini 3.7 Flash from Google.

How many models are evaluated on LVBench?

26 models have been evaluated on the LVBench benchmark, with 0 verified results and 26 self-reported results.

Where can I find the LVBench paper?

The LVBench paper is available at https://arxiv.org/abs/2406.08035. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does LVBench cover?

LVBench is categorized under long context, multimodal, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on LVBench?

Kimi K2.5 by Moonshot AI is the top-ranked open-source model on LVBench, with a score of 0.759 (rank #7).

Which model offers the best value on LVBench?

Among models scoring within 10% of the leader, Gemini 3.7 Flash from Google is the cheapest, at $0.75 per million input tokens with a score of 0.854.

How recent are the LVBench leaderboard results?

The LVBench leaderboard was last updated in August 2026 and currently includes 26 evaluated models.