NOVA-63
Progress Over Time
Interactive timeline showing model performance evolution on NOVA-63
State-of-the-art frontier
Open
Proprietary
NOVA-63 Leaderboard
11 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 397B | — | — | ||
| 2 | Alibaba Cloud / Qwen Team | — | 1.0M | $1.25 / $3.75 | ||
| 3 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 4 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 6 | Alibaba Cloud / Qwen Team | — | 1.0M | $0.50 / $3.00 | ||
| 7 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 8 | Alibaba Cloud / Qwen Team | 9B | — | — | ||
| 9 | Alibaba Cloud / Qwen Team | 4B | — | — | ||
| 10 | Alibaba Cloud / Qwen Team | 2B | — | — | ||
| 11 | Alibaba Cloud / Qwen Team | 800M | — | — |
Notice missing or incorrect data?
What is NOVA-63?
NOVA-63 is a multilingual evaluation benchmark covering 63 languages, designed to assess LLM performance across diverse linguistic contexts and tasks.
NOVA-63 is a text benchmark evaluating models on general tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for general leaderboards.
Current leaders
Qwen3.5-397B-A17B from Alibaba Cloud / Qwen Team currently leads the NOVA-63 leaderboard with a score of 0.591 across 11 evaluated AI models.
FAQ
Common questions about the NOVA-63 benchmark and leaderboard.