The AI arena is free today

Open Superagent

NOVA-63

Progress Over Time

Interactive timeline showing model performance evolution on NOVA-63

State-of-the-art frontier
Open
Proprietary

NOVA-63 Leaderboard

11 models
ContextCostLicense
1
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
397B
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$1.25 / $3.75
3
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
6
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
1.0M$0.50 / $3.00
7
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
8
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
9B
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
4B
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
2B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
800M
Notice missing or incorrect data?
About this benchmark

What is NOVA-63?

NOVA-63 is a multilingual evaluation benchmark covering 63 languages, designed to assess LLM performance across diverse linguistic contexts and tasks.

NOVA-63 is a text benchmark evaluating models on general tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.

Compare leaders on the best AI for general leaderboards.

Current leaders

Qwen3.5-397B-A17B from Alibaba Cloud / Qwen Team currently leads the NOVA-63 leaderboard with a score of 0.591 across 11 evaluated AI models.

1Qwen3.5-397B-A17BAlibaba Cloud / Qwen Team59.1%
2Qwen3.7 MaxAlibaba Cloud / Qwen Team59.0%
3Qwen3.7-PlusAlibaba Cloud / Qwen Team58.8%

FAQ

Common questions about the NOVA-63 benchmark and leaderboard.

What is the NOVA-63 benchmark?

NOVA-63 is a multilingual evaluation benchmark covering 63 languages, designed to assess LLM performance across diverse linguistic contexts and tasks.

What is the NOVA-63 leaderboard?

The NOVA-63 leaderboard ranks 11 AI models based on their performance on this benchmark. Currently, Qwen3.5-397B-A17B by Alibaba Cloud / Qwen Team leads with a score of 0.591. The average score across all models is 0.552.

What is the highest NOVA-63 score?

The highest NOVA-63 score is 0.591, achieved by Qwen3.5-397B-A17B from Alibaba Cloud / Qwen Team.

How many models are evaluated on NOVA-63?

11 models have been evaluated on the NOVA-63 benchmark, with 0 verified results and 11 self-reported results.

What categories does NOVA-63 cover?

NOVA-63 is categorized under general. The benchmark evaluates text models with multilingual support.

What is the best open-source model on NOVA-63?

Qwen3.5-397B-A17B by Alibaba Cloud / Qwen Team is the top-ranked open-source model on NOVA-63, with a score of 0.591 (rank #1).

Which model offers the best value on NOVA-63?

Among models scoring within 10% of the leader, Qwen3.5-27B from Alibaba Cloud / Qwen Team is the cheapest, at $0.30 per million input tokens with a score of 0.581.

How recent are the NOVA-63 leaderboard results?

The NOVA-63 leaderboard was last updated in August 2026 and currently includes 11 evaluated models.