WMT23
Progress Over Time
Interactive timeline showing model performance evolution on WMT23
WMT23 Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | — | — | ||
| 2 | Google | — | — | — | ||
| 3 | Google | 8B | — | — | ||
| 4 | Google | — | — | — |
What is WMT23?
The Eighth Conference on Machine Translation (WMT23) benchmark evaluating machine translation systems across 8 language pairs (14 translation directions) including general, biomedical, literary, and low-resource language translation tasks. Features specialized shared tasks for quality estimation, metrics evaluation, sign language translation, and discourse-level literary translation with professional human assessment.
WMT23 is a text benchmark evaluating models on language and healthcare tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 0.7, with the leader at 0.8.
Compare leaders on the best AI for language and best AI for healthcare leaderboards.
Current leaders
Gemini 1.5 Pro from Google currently leads the WMT23 leaderboard with a score of 0.751 across 4 evaluated AI models.
FAQ
Common questions about the WMT23 benchmark and leaderboard.