The AI arena is free today

Open Superagent

OmniGAIA

Progress Over Time

Interactive timeline showing model performance evolution on OmniGAIA

State-of-the-art frontier
Open
Proprietary

OmniGAIA Leaderboard

1 models
ContextCostLicense
1
Notice missing or incorrect data?
About this benchmark

What is OmniGAIA?

OmniGAIA evaluates multimodal perception and reasoning in agentic contexts, testing a model's ability to process diverse inputs and perform complex multi-step reasoning tasks.

OmniGAIA is a multimodal benchmark evaluating models on multimodal, reasoning, and agents tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–1 scale. The current average is 0.5, with the leader at 0.5.

Compare leaders on the best AI for multimodal, best AI for reasoning and best AI for agents leaderboards.

Current leaders

MiMo-V2-Omni from Xiaomi currently leads the OmniGAIA leaderboard with a score of 0.498 across 1 evaluated AI models.

1MiMo-V2-OmniXiaomi49.8%

FAQ

Common questions about the OmniGAIA benchmark and leaderboard.

What is the OmniGAIA benchmark?

OmniGAIA evaluates multimodal perception and reasoning in agentic contexts, testing a model's ability to process diverse inputs and perform complex multi-step reasoning tasks.

What is the OmniGAIA leaderboard?

The OmniGAIA leaderboard ranks 1 AI models based on their performance on this benchmark. Currently, MiMo-V2-Omni by Xiaomi leads with a score of 0.498. The average score across all models is 0.498.

What is the highest OmniGAIA score?

The highest OmniGAIA score is 0.498, achieved by MiMo-V2-Omni from Xiaomi.

How many models are evaluated on OmniGAIA?

1 models have been evaluated on the OmniGAIA benchmark, with 0 verified results and 1 self-reported results.

What categories does OmniGAIA cover?

OmniGAIA is categorized under multimodal, reasoning, and agents. The benchmark evaluates multimodal models.

How recent are the OmniGAIA leaderboard results?

The OmniGAIA leaderboard was last updated in August 2026 and currently includes 1 evaluated models.