HiddenMath
Progress Over Time
Interactive timeline showing model performance evolution on HiddenMath
HiddenMath Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | — | — | ||
| 2 | Google | 27B | — | — | ||
| 3 | Google | — | — | — | ||
| 4 | Google | 12B | — | — | ||
| 5 | Google | — | — | — | ||
| 6 | Google | — | — | — | ||
| 7 | Google | 4B | — | — | ||
| 8 | 2B | — | — | |||
| 8 | Google | 8B | — | — | ||
| 10 | Google | 8B | — | — | ||
| 11 | Google | 8B | — | — | ||
| 11 | 2B | — | — | |||
| 13 | Google | 1B | — | — |
What is HiddenMath?
Google DeepMind's internal mathematical reasoning benchmark that introduces novel problems not encountered during model training to evaluate true mathematical reasoning capabilities rather than memorization
HiddenMath is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 13 models on this benchmark, scored on a 0–1 scale. The current average is 0.4, with the leader at 0.6.
Compare leaders on the best AI for math and best AI for reasoning leaderboards.
Current leaders
Gemini 2.0 Flash from Google currently leads the HiddenMath leaderboard with a score of 0.630 across 13 evaluated AI models.
FAQ
Common questions about the HiddenMath benchmark and leaderboard.