OSWorld-G
Progress Over Time
Interactive timeline showing model performance evolution on OSWorld-G
OSWorld-G Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Alibaba Cloud / Qwen Team | 236B | — | — |
What is OSWorld-G?
OSWorld-G (Grounding) evaluates screenshot grounding accuracy for OS automation tasks.
OSWorld-G is a image benchmark evaluating models on multimodal, grounding, agents, and vision tasks. LLM Stats tracks 1 models on this benchmark, scored on a 0–100 scale. The current average is 0.7, with the leader at 0.7.
Compare leaders on the best AI for multimodal, best AI for grounding, best AI for agents and best AI for vision leaderboards.
Current leaders
Qwen3 VL 235B A22B Thinking from Alibaba Cloud / Qwen Team currently leads the OSWorld-G leaderboard with a score of 0.683 across 1 evaluated AI models.
FAQ
Common questions about the OSWorld-G benchmark and leaderboard.