The AI arena is free today

Open Superagent

PhysicsFinals

Paper

Progress Over Time

Interactive timeline showing model performance evolution on PhysicsFinals

State-of-the-art frontier
Open
Proprietary

PhysicsFinals Leaderboard

2 models
ContextCostLicense
1
2
Notice missing or incorrect data?
About this benchmark

What is PhysicsFinals?

PHYSICS is a comprehensive benchmark for university-level physics problem solving, containing 1,297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. Even advanced models like o3-mini achieve only 59.9% accuracy.

PhysicsFinals is a text benchmark evaluating models on math, physics, and reasoning tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.

Compare leaders on the best AI for math, best AI for physics and best AI for reasoning leaderboards.

Current leaders

Gemini 1.5 Pro from Google currently leads the PhysicsFinals leaderboard with a score of 0.639 across 2 evaluated AI models.

1Gemini 1.5 ProGoogle63.9%
2Gemini 1.5 FlashGoogle57.4%

Source paper

Title
PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
Authors
Kaiyue Feng, Yilun Zhao, Yixin Liu, Tianyu Yang, and 3 others
Published
Abstract

We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. We develop a robust automated evaluation system for precise and reliable validation. Our evaluation of leading foundation models reveals substantial limitations. Even the most advanced model, o3-mini, achieves only 59.9% accuracy, highlighting significant challenges in solving high-level scientific problems. Through comprehensive error analysis, exploration of diverse prompting strategies, and Retrieval-Augmented Generation (RAG)-based knowledge augmentation, we identify key areas for improvement, laying the foundation for future advancements.

FAQ

Common questions about the PhysicsFinals benchmark and leaderboard.

What is the PhysicsFinals benchmark?

PHYSICS is a comprehensive benchmark for university-level physics problem solving, containing 1,297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. Even advanced models like o3-mini achieve only 59.9% accuracy.

What is the PhysicsFinals leaderboard?

The PhysicsFinals leaderboard ranks 2 AI models based on their performance on this benchmark. Currently, Gemini 1.5 Pro by Google leads with a score of 0.639. The average score across all models is 0.607.

What is the highest PhysicsFinals score?

The highest PhysicsFinals score is 0.639, achieved by Gemini 1.5 Pro from Google.

How many models are evaluated on PhysicsFinals?

2 models have been evaluated on the PhysicsFinals benchmark, with 0 verified results and 2 self-reported results.

Where can I find the PhysicsFinals paper?

The PhysicsFinals paper is available at https://arxiv.org/abs/2503.21821. The paper details the methodology, dataset construction, and evaluation criteria.

What categories does PhysicsFinals cover?

PhysicsFinals is categorized under math, physics, and reasoning. The benchmark evaluates text models.

How recent are the PhysicsFinals leaderboard results?

The PhysicsFinals leaderboard was last updated in August 2026 and currently includes 2 evaluated models.