The AI arena is free today

Open Superagent

CoVoST2

Paper

Progress Over Time

Interactive timeline showing model performance evolution on CoVoST2

State-of-the-art frontier
Open
Proprietary

CoVoST2 Leaderboard

4 models
ContextCostLicense
1
2
312B
4
Notice missing or incorrect data?
About this benchmark

What is CoVoST2?

CoVoST 2 is a large-scale multilingual speech translation corpus derived from Common Voice, covering translations from 21 languages into English and from English into 15 languages. The dataset contains 2,880 hours of speech with 78K speakers for speech translation research.

CoVoST2 is a audio benchmark evaluating models on language, speech to text, and audio tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.4.

Compare leaders on the best AI for language, best AI for speech to text and best AI for audio leaderboards.

Current leaders

Nova 2 Omni from Amazon currently leads the CoVoST2 leaderboard with a score of 0.407 across 4 evaluated AI models.

1Nova 2 OmniAmazon40.7%
2Gemini 2.0 FlashGoogle39.2%
3Gemma 4 12BGoogle38.5%

Source paper

Title
CoVoST 2 and Massively Multilingual Speech-to-Text Translation
Authors
Changhan Wang, Anne Wu, Juan Pino
Published
Abstract

Speech translation has recently become an increasingly popular topic of research, partly due to the development of benchmark datasets. Nevertheless, current datasets cover a limited number of languages. With the aim to foster research in massive multilingual speech translation and speech translation for low resource language pairs, we release CoVoST 2, a large-scale multilingual speech translation corpus covering translations from 21 languages into English and from English into 15 languages. This represents the largest open dataset available to date from total volume and language coverage perspective. Data sanity checks provide evidence about the quality of the data, which is released under CC0 license. We also provide extensive speech recognition, bilingual and multilingual machine translation and speech translation baselines with open-source implementation.

FAQ

Common questions about the CoVoST2 benchmark and leaderboard.

What is the CoVoST2 benchmark?

CoVoST 2 is a large-scale multilingual speech translation corpus derived from Common Voice, covering translations from 21 languages into English and from English into 15 languages. The dataset contains 2,880 hours of speech with 78K speakers for speech translation research.

What is the CoVoST2 leaderboard?

The CoVoST2 leaderboard ranks 4 AI models based on their performance on this benchmark. Currently, Nova 2 Omni by Amazon leads with a score of 0.407. The average score across all models is 0.392.

What is the highest CoVoST2 score?

The highest CoVoST2 score is 0.407, achieved by Nova 2 Omni from Amazon.

How many models are evaluated on CoVoST2?

4 models have been evaluated on the CoVoST2 benchmark, with 0 verified results and 4 self-reported results.

Where can I find the CoVoST2 paper?

The CoVoST2 paper is available at https://arxiv.org/abs/2007.10310. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does CoVoST2 cover?

CoVoST2 is categorized under language, speech to text, and audio. The benchmark evaluates audio models with multilingual support.

What is the best open-source model on CoVoST2?

Gemma 4 12B by Google is the top-ranked open-source model on CoVoST2, with a score of 0.385 (rank #3).

How recent are the CoVoST2 leaderboard results?

The CoVoST2 leaderboard was last updated in August 2026 and currently includes 4 evaluated models.