BabyVision
Progress Over Time
Interactive timeline showing model performance evolution on BabyVision
State-of-the-art frontier
Open
Proprietary
BabyVision Leaderboard
11 models
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | Alibaba Cloud / Qwen Team | 28B | — | — | ||
| 3 | Meta | — | 1.0M | $1.25 / $4.25 | ||
| 4 | ByteDance | — | — | — | ||
| 5 | Alibaba Cloud / Qwen Team | — | — | — | ||
| 6 | Moonshot AI | 1.0T | 262K | $0.75 / $3.50 | ||
| 7 | ByteDance | — | — | — | ||
| 8 | Zhipu AI | 320B | 1.0M | $0.15 / $0.50 | ||
| 9 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 10 | Alibaba Cloud / Qwen Team | 122B | — | — | ||
| 11 | Alibaba Cloud / Qwen Team | 35B | — | — |
Notice missing or incorrect data?
What is BabyVision?
A benchmark for early-stage visual reasoning and perception on child-like vision tasks.
BabyVision is a multimodal benchmark evaluating models on multimodal, reasoning, and vision tasks. LLM Stats tracks 11 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.9.
Compare leaders on the best AI for multimodal, best AI for reasoning and best AI for vision leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the BabyVision leaderboard with a score of 0.857 across 11 evaluated AI models.
FAQ
Common questions about the BabyVision benchmark and leaderboard.