The AI arena is free today

Open Superagent

Translation Set1→en spBleu

Paper

Progress Over Time

Interactive timeline showing model performance evolution on Translation Set1→en spBleu

State-of-the-art frontier
Open
Proprietary

Translation Set1→en spBleu Leaderboard

3 models
ContextCostLicense
1
Amazon
Amazon
2
Amazon
Amazon
3
Notice missing or incorrect data?
About this benchmark

What is Translation Set1→en spBleu?

spBLEU (SentencePiece BLEU) evaluation metric for machine translation quality assessment, using language-agnostic SentencePiece tokenization with BLEU scoring. Part of the FLORES-101 evaluation benchmark for low-resource and multilingual machine translation.

Translation Set1→en spBleu is a text benchmark evaluating models on language tasks. LLM Stats tracks 3 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.4.

Compare leaders on the best AI for language leaderboards.

Current leaders

Nova Pro from Amazon currently leads the Translation Set1→en spBleu leaderboard with a score of 0.444 across 3 evaluated AI models.

1Nova ProAmazon44.4%
2Nova LiteAmazon43.1%
3Nova MicroAmazon42.6%

Source paper

Title
The FLORES-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation
Authors
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, and 6 others
Published
Abstract

One of the biggest challenges hindering progress in low-resource and multilingual machine translation is the lack of good evaluation benchmarks. Current evaluation benchmarks either lack good coverage of low-resource languages, consider only restricted domains, or are low quality because they are constructed using semi-automatic procedures. In this work, we introduce the FLORES-101 evaluation benchmark, consisting of 3001 sentences extracted from English Wikipedia and covering a variety of different topics and domains. These sentences have been translated in 101 languages by professional translators through a carefully controlled process. The resulting dataset enables better assessment of model quality on the long tail of low-resource languages, including the evaluation of many-to-many multilingual translation systems, as all translations are multilingually aligned. By publicly releasing such a high-quality and high-coverage dataset, we hope to foster progress in the machine translation community and beyond.

FAQ

Common questions about the Translation Set1→en spBleu benchmark and leaderboard.

What is the Translation Set1→en spBleu benchmark?

spBLEU (SentencePiece BLEU) evaluation metric for machine translation quality assessment, using language-agnostic SentencePiece tokenization with BLEU scoring. Part of the FLORES-101 evaluation benchmark for low-resource and multilingual machine translation.

What is the Translation Set1→en spBleu leaderboard?

The Translation Set1→en spBleu leaderboard ranks 3 AI models based on their performance on this benchmark. Currently, Nova Pro by Amazon leads with a score of 0.444. The average score across all models is 0.434.

What is the highest Translation Set1→en spBleu score?

The highest Translation Set1→en spBleu score is 0.444, achieved by Nova Pro from Amazon.

How many models are evaluated on Translation Set1→en spBleu?

3 models have been evaluated on the Translation Set1→en spBleu benchmark, with 0 verified results and 3 self-reported results.

Where can I find the Translation Set1→en spBleu paper?

The Translation Set1→en spBleu paper is available at https://arxiv.org/abs/2106.03193. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does Translation Set1→en spBleu cover?

Translation Set1→en spBleu is categorized under language. The benchmark evaluates text models with multilingual support.

How recent are the Translation Set1→en spBleu leaderboard results?

The Translation Set1→en spBleu leaderboard was last updated in August 2026 and currently includes 3 evaluated models.