The AI arena is free today

Open Superagent

GAIA2

Progress Over Time

Interactive timeline showing model performance evolution on GAIA2

State-of-the-art frontier
Open
Proprietary

GAIA2 Leaderboard

1 models
ContextCostLicense
130B
Notice missing or incorrect data?
About this benchmark

What is GAIA2?

GAIA2 evaluates general-purpose AI agents on real-world, multi-step questions that require reasoning, tool use, and information retrieval.

GAIA2 is a multimodal benchmark evaluating models on reasoning, general, agents, and tool calling tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.4.

Compare leaders on the best AI for reasoning, best AI for general, best AI for agents and best AI for tool calling leaderboards.

Current leaders

Muse Glimmer-30B from Meta currently leads the GAIA2 leaderboard with a score of 0.433 across 1 evaluated AI models.

FAQ

Common questions about the GAIA2 benchmark and leaderboard.

What is the GAIA2 benchmark?

GAIA2 evaluates general-purpose AI agents on real-world, multi-step questions that require reasoning, tool use, and information retrieval.

What is the GAIA2 leaderboard?

The GAIA2 leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, Muse Glimmer-30B by Meta leads with a score of 0.433. The average score across all models is 0.433.

What is the highest GAIA2 score?

The highest GAIA2 score is 0.433, achieved by Muse Glimmer-30B from Meta.

How many models are evaluated on GAIA2?

1 models have been evaluated on the GAIA2 benchmark, with 0 verified results and 1 self-reported results.

What categories does GAIA2 cover?

GAIA2 is categorized under reasoning, general, agents, and tool calling. The benchmark evaluates multimodal models.

What is the best open-source model on GAIA2?

Muse Glimmer-30B by Meta is the top-ranked open-source model on GAIA2, with a score of 0.433 (rank #1).

How recent are the GAIA2 leaderboard results?

The GAIA2 leaderboard was last updated in August 2026 and currently includes 1 evaluated models.