The AI arena is free today

Open Superagent

FActScore

Paper

Progress Over Time

Interactive timeline showing model performance evolution on FActScore

State-of-the-art frontier
Open
Proprietary

FActScore Leaderboard

3 models
ContextCostLicense
1
2
ByteDance
ByteDance
256K$0.10 / $0.40
3
OpenAI
OpenAI
Notice missing or incorrect data?
About this benchmark

What is FActScore?

A fine-grained atomic evaluation metric for factual precision in long-form text generation that breaks generated text into atomic facts and computes the percentage supported by reliable knowledge sources, with automated assessment using retrieval and language models

FActScore is a text benchmark evaluating models on reasoning tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 1.0.

Compare leaders on the best AI for reasoning leaderboards.

Current leaders

Grok-4.1 from xAI currently leads the FActScore leaderboard with a score of 0.970 across 3 evaluated AI models.

1Grok-4.1xAI97.0%
2Seed 2.0 MiniByteDance50.4%
3GPT-5OpenAI1.0%

FAQ

Common questions about the FActScore benchmark and leaderboard.

What is the FActScore benchmark?

A fine-grained atomic evaluation metric for factual precision in long-form text generation that breaks generated text into atomic facts and computes the percentage supported by reliable knowledge sources, with automated assessment using retrieval and language models

What is the FActScore leaderboard?

The FActScore leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Grok-4.1 by xAI leads with a score of 0.970. The average score across all models is 0.495.

What is the highest FActScore score?

The highest FActScore score is 0.970, achieved by Grok-4.1 from xAI.

How many models are evaluated on FActScore?

3 models have been evaluated on the FActScore benchmark, with 0 verified results and 3 self-reported results.

Where can I find the FActScore paper?

The FActScore paper is available at https://arxiv.org/abs/2305.14251. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does FActScore cover?

FActScore is categorized under reasoning. The benchmark evaluates text models.

How recent are the FActScore leaderboard results?

The FActScore leaderboard was last updated in September 2026 and currently includes 3 evaluated models.