The AI arena is free today

Open Superagent

InfiniteBench/En.QA

Paper

Progress Over Time

Interactive timeline showing model performance evolution on InfiniteBench/En.QA

State-of-the-art frontier
Open
Proprietary

InfiniteBench/En.QA Leaderboard

1 models
ContextCostLicense
13B
Notice missing or incorrect data?
About this benchmark

What is InfiniteBench/En.QA?

InfiniteBench English Question Answering variant - first LLM benchmark featuring average data length surpassing 100K tokens for evaluating long-context capabilities with 12 tasks spanning diverse domains

InfiniteBench/En.QA is a text benchmark evaluating models on long context tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.2, with the leader at 0.2.

Compare leaders on the best AI for long context leaderboards.

Current leaders

Llama 3.2 3B Instruct from Meta currently leads the InfiniteBench/En.QA leaderboard with a score of 0.198 across 1 evaluated AI models.

Source paper

Title
$\infty$Bench: Extending Long Context Evaluation Beyond 100K Tokens
Authors
Xinrong Zhang, Yingfa Chen, Shengding Hu, Zihang Xu, and 7 others
Published
Abstract

Processing and reasoning over long contexts is crucial for many practical applications of Large Language Models (LLMs), such as document comprehension and agent construction. Despite recent strides in making LLMs process contexts with more than 100K tokens, there is currently a lack of a standardized benchmark to evaluate this long-context capability. Existing public benchmarks typically focus on contexts around 10K tokens, limiting the assessment and comparison of LLMs in processing longer contexts. In this paper, we propose $\infty$Bench, the first LLM benchmark featuring an average data length surpassing 100K tokens. $\infty$Bench comprises synthetic and realistic tasks spanning diverse domains, presented in both English and Chinese. The tasks in $\infty$Bench are designed to require well understanding of long dependencies in contexts, and make simply retrieving a limited number of passages from contexts not sufficient for these tasks. In our experiments, based on $\infty$Bench, we evaluate the state-of-the-art proprietary and open-source LLMs tailored for processing long contexts. The results indicate that existing long context LLMs still require significant advancements to effectively process 100K+ context. We further present three intriguing analyses regarding the behavior of LLMs processing long context.

FAQ

Common questions about the InfiniteBench/En.QA benchmark and leaderboard.

What is the InfiniteBench/En.QA benchmark?

InfiniteBench English Question Answering variant - first LLM benchmark featuring average data length surpassing 100K tokens for evaluating long-context capabilities with 12 tasks spanning diverse domains

What is the InfiniteBench/En.QA leaderboard?

The InfiniteBench/En.QA leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Llama 3.2 3B Instruct by Meta leads with a score of 0.198. The average score across all models is 0.198.

What is the highest InfiniteBench/En.QA score?

The highest InfiniteBench/En.QA score is 0.198, achieved by Llama 3.2 3B Instruct from Meta.

How many models are evaluated on InfiniteBench/En.QA?

1 models have been evaluated on the InfiniteBench/En.QA benchmark, with 0 verified results and 1 self-reported results.

Where can I find the InfiniteBench/En.QA paper?

The InfiniteBench/En.QA paper is available at https://arxiv.org/abs/2402.13718. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does InfiniteBench/En.QA cover?

InfiniteBench/En.QA is categorized under long context. The benchmark evaluates text models.

How recent are the InfiniteBench/En.QA leaderboard results?

The InfiniteBench/En.QA leaderboard was last updated in September 2026 and currently includes 1 evaluated models.