The AI arena is free today

Open Superagent

BabyVision

Progress Over Time

Interactive timeline showing model performance evolution on BabyVision

State-of-the-art frontier
Open
Proprietary

BabyVision Leaderboard

11 models
ContextCostLicense
1
Moonshot AI
Moonshot AI
2.8T1.0M$3.00 / $15.00
2
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
28B
31.0M$1.25 / $4.25
4
ByteDance
ByteDance
5
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
6
Moonshot AI
Moonshot AI
1.0T262K$0.75 / $3.50
7
8
Zhipu AI
Zhipu AI
320B1.0M$0.15 / $0.50
9
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
27B262K$0.30 / $2.40
10
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
122B
11
Alibaba Cloud / Qwen Team
Alibaba Cloud / Qwen Team
35B
Notice missing or incorrect data?
About this benchmark

What is BabyVision?

A benchmark for early-stage visual reasoning and perception on child-like vision tasks.

BabyVision is a multimodal benchmark evaluating models on multimodal, reasoning, and vision tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.9.

Compare leaders on the best AI for multimodal, best AI for reasoning and best AI for vision leaderboards.

Current leaders

Kimi K3 from Moonshot AI currently leads the BabyVision leaderboard with a score of 0.857 across 11 evaluated AI models.

1Kimi K3Moonshot AI85.7%
2Qwen3.8-27BAlibaba Cloud / Qwen Team85.6%
3Muse Spark 1.1Meta76.3%

FAQ

Common questions about the BabyVision benchmark and leaderboard.

What is the BabyVision benchmark?

A benchmark for early-stage visual reasoning and perception on child-like vision tasks.

What is the BabyVision leaderboard?

The BabyVision leaderboard ranks 11 AI models based on their performance on this benchmark. Currently, Kimi K3 by Moonshot AI leads with a score of 0.857. The average score across all models is 0.636.

What is the highest BabyVision score?

The highest BabyVision score is 0.857, achieved by Kimi K3 from Moonshot AI.

How many models are evaluated on BabyVision?

11 models have been evaluated on the BabyVision benchmark, with 0 verified results and 11 self-reported results.

What categories does BabyVision cover?

BabyVision is categorized under multimodal, reasoning, and vision. The benchmark evaluates multimodal models.

What is the best open-source model on BabyVision?

Kimi K3 by Moonshot AI is the top-ranked open-source model on BabyVision, with a score of 0.857 (rank #1).

Which model offers the best value on BabyVision?

Among models scoring within 10% of the leader, Kimi K3 from Moonshot AI is the cheapest, at $3.00 per million input tokens with a score of 0.857.

How recent are the BabyVision leaderboard results?

The BabyVision leaderboard was last updated in August 2026 and currently includes 11 evaluated models.