OpenAI-MRCR: 2 needle 256k

Paper

Progress Over Time

Interactive timeline showing model performance evolution on OpenAI-MRCR: 2 needle 256k

State-of-the-art frontier
Open
Proprietary

OpenAI-MRCR: 2 needle 256k Leaderboard

1 models
ContextCostLicense
1
OpenAI
OpenAI
Notice missing or incorrect data?
About this benchmark

What is OpenAI-MRCR: 2 needle 256k?

Multi-Round Co-reference Resolution (MRCR) benchmark that tests long-context reasoning by evaluating a model's ability to distinguish between similar outputs, reason about ordering, and reproduce specific content from multi-turn conversations containing multiple writing requests on overlapping topics at 256k tokens.

OpenAI-MRCR: 2 needle 256k is a text benchmark evaluating models on reasoning and long context tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.9, with the leader at 0.9.

Compare leaders on the best AI for reasoning and best AI for long context leaderboards.

Current leaders

GPT-5 from OpenAI currently leads the OpenAI-MRCR: 2 needle 256k leaderboard with a score of 0.868 across 1 evaluated AI models.

1GPT-5OpenAI86.8%

Source paper

Title
Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
Authors
Kiran Vodrahalli, Santiago Ontanon, Nilesh Tripuraneni, Kelvin Xu, and 20 others
Published
Abstract

We introduce Michelangelo: a minimal, synthetic, and unleaked long-context reasoning evaluation for large language models which is also easy to automatically score. This evaluation is derived via a novel, unifying framework for evaluations over arbitrarily long contexts which measure the model's ability to do more than retrieve a single piece of information from its context. The central idea of the Latent Structure Queries framework (LSQ) is to construct tasks which require a model to ``chisel away'' the irrelevant information in the context, revealing a latent structure in the context. To verify a model's understanding of this latent structure, we query the model for details of the structure. Using LSQ, we produce three diagnostic long-context evaluations across code and natural-language domains intended to provide a stronger signal of long-context language model capabilities. We perform evaluations on several state-of-the-art models and demonstrate both that a) the proposed evaluations are high-signal and b) that there is significant room for improvement in synthesizing long-context information.

FAQ

Common questions about the OpenAI-MRCR: 2 needle 256k benchmark and leaderboard.

What is the OpenAI-MRCR: 2 needle 256k benchmark?

Multi-Round Co-reference Resolution (MRCR) benchmark that tests long-context reasoning by evaluating a model's ability to distinguish between similar outputs, reason about ordering, and reproduce specific content from multi-turn conversations containing multiple writing requests on overlapping topics at 256k tokens.

What is the OpenAI-MRCR: 2 needle 256k leaderboard?

The OpenAI-MRCR: 2 needle 256k leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, GPT-5 by OpenAI leads with a score of 0.868. The average score across all models is 0.868.

What is the highest OpenAI-MRCR: 2 needle 256k score?

The highest OpenAI-MRCR: 2 needle 256k score is 0.868, achieved by GPT-5 from OpenAI.

How many models are evaluated on OpenAI-MRCR: 2 needle 256k?

1 models have been evaluated on the OpenAI-MRCR: 2 needle 256k benchmark, with 0 verified results and 1 self-reported results.

Where can I find the OpenAI-MRCR: 2 needle 256k paper?

The OpenAI-MRCR: 2 needle 256k paper is available at https://arxiv.org/abs/2409.12640. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does OpenAI-MRCR: 2 needle 256k cover?

OpenAI-MRCR: 2 needle 256k is categorized under reasoning and long context. The benchmark evaluates text models.

How recent are the OpenAI-MRCR: 2 needle 256k leaderboard results?

The OpenAI-MRCR: 2 needle 256k leaderboard was last updated in July 2026 and currently includes 1 evaluated models.