The AI arena is free today

Open Superagent

FunctionalMATH

Paper

Progress Over Time

Interactive timeline showing model performance evolution on FunctionalMATH

State-of-the-art frontier
Open
Proprietary

FunctionalMATH Leaderboard

2 models
ContextCostLicense
1
2
Notice missing or incorrect data?
About this benchmark

What is FunctionalMATH?

A functional variant of the MATH benchmark that tests language models' ability to generalize reasoning patterns across different problem instances, revealing the reasoning gap between static and functional performance.

FunctionalMATH is a text benchmark evaluating models on math and reasoning tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.

Compare leaders on the best AI for math and best AI for reasoning leaderboards.

Current leaders

Gemini 1.5 Pro from Google currently leads the FunctionalMATH leaderboard with a score of 0.646 across 2 evaluated AI models.

1Gemini 1.5 ProGoogle64.6%
2Gemini 1.5 FlashGoogle53.6%

FAQ

Common questions about the FunctionalMATH benchmark and leaderboard.

What is the FunctionalMATH benchmark?

A functional variant of the MATH benchmark that tests language models' ability to generalize reasoning patterns across different problem instances, revealing the reasoning gap between static and functional performance.

What is the FunctionalMATH leaderboard?

The FunctionalMATH leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Gemini 1.5 Pro by Google leads with a score of 0.646. The average score across all models is 0.591.

What is the highest FunctionalMATH score?

The highest FunctionalMATH score is 0.646, achieved by Gemini 1.5 Pro from Google.

How many models are evaluated on FunctionalMATH?

2 models have been evaluated on the FunctionalMATH benchmark, with 0 verified results and 2 self-reported results.

Where can I find the FunctionalMATH paper?

The FunctionalMATH paper is available at https://arxiv.org/abs/2402.19450. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does FunctionalMATH cover?

FunctionalMATH is categorized under math and reasoning. The benchmark evaluates text models.

How recent are the FunctionalMATH leaderboard results?

The FunctionalMATH leaderboard was last updated in August 2026 and currently includes 2 evaluated models.