VLMsAreBlind
Progress Over Time
Interactive timeline showing model performance evolution on VLMsAreBlind
VLMsAreBlind Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 35B | — | — | ||
| 1 | Alibaba Cloud / Qwen Team | 28B | 262K | $0.60 / $3.60 | ||
| 3 | Alibaba Cloud / Qwen Team | 27B | 262K | $0.30 / $2.40 | ||
| 4 | Alibaba Cloud / Qwen Team | 122B | — | — |
What is VLMsAreBlind?
A vision-language benchmark that probes blind spots and brittle reasoning in multimodal models.
VLMsAreBlind is a multimodal benchmark evaluating models on multimodal, reasoning, and vision tasks. LLM Stats tracks 4 models on this benchmark, scored on a 0–1 scale. The current average is 1.0, with the leader at 1.0.
Compare leaders on the best AI for multimodal, best AI for reasoning and best AI for vision leaderboards.
Current leaders
Qwen3.5-35B-A3B from Alibaba Cloud / Qwen Team currently leads the VLMsAreBlind leaderboard with a score of 0.970 across 4 evaluated AI models.
FAQ
Common questions about the VLMsAreBlind benchmark and leaderboard.