FunctionalMATH
Progress Over Time
Interactive timeline showing model performance evolution on FunctionalMATH
FunctionalMATH Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | — | — | ||
| 2 | Google | — | — | — |
What is FunctionalMATH?
A functional variant of the MATH benchmark that tests language models' ability to generalize reasoning patterns across different problem instances, revealing the reasoning gap between static and functional performance.
FunctionalMATH is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for math and best AI for reasoning leaderboards.
Current leaders
Gemini 1.5 Pro from Google currently leads the FunctionalMATH leaderboard with a score of 0.646 across 2 evaluated AI models.
FAQ
Common questions about the FunctionalMATH benchmark and leaderboard.