MLS-Bench Lite
Progress Over Time
Interactive timeline showing model performance evolution on MLS-Bench Lite
MLS-Bench Lite Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Kimi K3New Moonshot AI | 2.8T | 1.0M | $3.00 / $15.00 | ||
| 2 | Moonshot AI | 1.0T | 262K | $0.74 / $3.50 |
What is MLS-Bench Lite?
MLS-Bench Lite is the official 30-task subset of MLS-Bench for evaluating whether AI systems can invent generalizable and scalable machine learning methods across LLM pretraining and post-training, robotics, world models, computer vision, reinforcement learning, optimization, ML systems, and AI for Science.
MLS-Bench Lite is a text benchmark evaluating models on reasoning, agents, and coding tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.5.
Compare leaders on the best AI for reasoning, best AI for agents and best AI for coding leaderboards.
Current leaders
Kimi K3 from Moonshot AI currently leads the MLS-Bench Lite leaderboard with a score of 0.483 across 2 evaluated AI models.
FAQ
Common questions about the MLS-Bench Lite benchmark and leaderboard.