LVBench
Progress Over Time
Interactive timeline showing model performance evolution on LVBench
LVBench Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | 1.0M | $0.75 / $3.75 | ||
| 2 | Alibaba Cloud / Qwen Team | 2.4T | 1.0M | $1.65 / $4.95 | ||
| 3 | ByteDance | — | — | — | ||
| 4 | ByteDance | — | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 125B | — | — | ||
| 6 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 7 | Moonshot AI | 1.0T | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 9 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 10 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 10 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 12 | Alibaba Cloud / Qwen Team | 236B | — | — | ||
| 13 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 14 | Alibaba Cloud / Qwen Team | 236B | — | — | ||
| 15 | Alibaba Cloud / Qwen Team | 33B | — | — | ||
| 16 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 17 | Alibaba Cloud / Qwen Team | 31B | — | — | ||
| 18 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 19 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $0.60 | ||
| 20 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 21 | Alibaba Cloud / Qwen Team | 4B | 262K | $0.10 / $1.00 | ||
| 22 | Alibaba Cloud / Qwen Team | 34B | — | — | ||
| 23 | Alibaba Cloud / Qwen Team | 72B | — | — | ||
| 24 | Alibaba Cloud / Qwen Team | 8B | — | — | ||
| 25 | Amazon | — | — | — | ||
| 26 | Amazon | — | — | — |
What is LVBench?
LVBench is an extreme long video understanding benchmark designed to evaluate multimodal models on videos up to two hours in duration. It contains 6 major categories and 21 subcategories, with videos averaging five times longer than existing datasets. The benchmark addresses applications requiring comprehension of extremely long videos.
LVBench is a multimodal benchmark evaluating models on long context, multimodal, and vision tasks. LLM Stats tracks 26 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.9.
Compare leaders on the best AI for long context, best AI for multimodal and best AI for vision leaderboards.
Current leaders
Gemini 3.7 Flash from Google currently leads the LVBench leaderboard with a score of 0.854 across 26 evaluated AI models.
Source paper
- Title
- LVBench: An Extreme Long Video Understanding Benchmark
- Authors
- Weihan Wang, Zehai He, Wenyi Hong, Yean Cheng, and 8 others
- Published
- arXiv
- 2406.08035
Abstract
Recent progress in multimodal large language models has markedly enhanced the understanding of short videos (typically under one minute), and several evaluation datasets have emerged accordingly. However, these advancements fall short of meeting the demands of real-world applications such as embodied intelligence for long-term decision-making, in-depth movie reviews and discussions, and live sports commentary, all of which require comprehension of long videos spanning several hours. To address this gap, we introduce LVBench, a benchmark specifically designed for long video understanding. Our dataset comprises publicly sourced videos and encompasses a diverse set of tasks aimed at long video comprehension and information extraction. LVBench is designed to challenge multimodal models to demonstrate long-term memory and extended comprehension capabilities. Our extensive evaluations reveal that current multimodal models still underperform on these demanding long video understanding tasks. Through LVBench, we aim to spur the development of more advanced models capable of tackling the complexities of long video comprehension. Our data and code are publicly available at: https://lvbench.github.io.
FAQ
Common questions about the LVBench benchmark and leaderboard.