ARC-E
Progress Over Time
Interactive timeline showing model performance evolution on ARC-E
ARC-E Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | 27B | — | — | ||
| 2 | Google | 9B | — | — | ||
| 3 | Nous Research | 70B | — | — | ||
| 4 | Google | 8B | — | — | ||
| 4 | 2B | — | — | |||
| 6 | 2B | — | — | |||
| 6 | Google | 8B | — | — | ||
| 8 | Baidu | 21B | — | — |
What is ARC-E?
ARC-E (AI2 Reasoning Challenge - Easy Set) is a subset of grade-school level, multiple-choice science questions that requires knowledge and reasoning capabilities. Part of the AI2 Reasoning Challenge dataset containing 5,197 questions that test scientific reasoning and factual knowledge. The Easy Set contains questions that are answerable by retrieval-based and word co-occurrence algorithms, making them more accessible than the Challenge Set.
ARC-E is a text benchmark evaluating models on reasoning and general tasks. LLM Stats tracks 8 models on this benchmark, scored on a 0–1 scale. The current average is 0.8, with the leader at 0.9.
Compare leaders on the best AI for reasoning and best AI for general leaderboards.
Current leaders
Gemma 2 27B from Google currently leads the ARC-E leaderboard with a score of 0.886 across 8 evaluated AI models.
Source paper
- Title
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Authors
- Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, and 3 others
- Published
- arXiv
- 1803.05457
Abstract
We present a new question set, text corpus, and baselines assembled to encourage AI research in advanced question answering. Together, these constitute the AI2 Reasoning Challenge (ARC), which requires far more powerful knowledge and reasoning than previous challenges such as SQuAD or SNLI. The ARC question set is partitioned into a Challenge Set and an Easy Set, where the Challenge Set contains only questions answered incorrectly by both a retrieval-based algorithm and a word co-occurence algorithm. The dataset contains only natural, grade-school science questions (authored for human tests), and is the largest public-domain set of this kind (7,787 questions). We test several baselines on the Challenge Set, including leading neural models from the SQuAD and SNLI tasks, and find that none are able to significantly outperform a random baseline, reflecting the difficult nature of this task. We are also releasing the ARC Corpus, a corpus of 14M science sentences relevant to the task, and implementations of the three neural baseline models tested. Can your model perform better? We pose ARC as a challenge to the community.
FAQ
Common questions about the ARC-E benchmark and leaderboard.