PhysicsFinals
Progress Over Time
Interactive timeline showing model performance evolution on PhysicsFinals
PhysicsFinals Leaderboard
| Context | Cost | License | ||||
|---|---|---|---|---|---|---|
| 1 | Google | — | — | — | ||
| 2 | Google | — | — | — |
What is PhysicsFinals?
PHYSICS is a comprehensive benchmark for university-level physics problem solving, containing 1,297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. Even advanced models like o3-mini achieve only 59.9% accuracy.
PhysicsFinals is a text benchmark evaluating models on math, physics, and reasoning tasks. LLM Stats tracks 2 models on this benchmark, scored on a 0–1 scale. The current average is 0.6, with the leader at 0.6.
Compare leaders on the best AI for math, best AI for physics and best AI for reasoning leaderboards.
Current leaders
Gemini 1.5 Pro from Google currently leads the PhysicsFinals leaderboard with a score of 0.639 across 2 evaluated AI models.
Source paper
- Title
- PHYSICS: Benchmarking Foundation Models on University-Level Physics Problem Solving
- Authors
- Kaiyue Feng, Yilun Zhao, Yixin Liu, Tianyu Yang, and 3 others
- Published
- arXiv
- 2503.21821
Abstract
We introduce PHYSICS, a comprehensive benchmark for university-level physics problem solving. It contains 1297 expert-annotated problems covering six core areas: classical mechanics, quantum mechanics, thermodynamics and statistical mechanics, electromagnetism, atomic physics, and optics. Each problem requires advanced physics knowledge and mathematical reasoning. We develop a robust automated evaluation system for precise and reliable validation. Our evaluation of leading foundation models reveals substantial limitations. Even the most advanced model, o3-mini, achieves only 59.9% accuracy, highlighting significant challenges in solving high-level scientific problems. Through comprehensive error analysis, exploration of diverse prompting strategies, and Retrieval-Augmented Generation (RAG)-based knowledge augmentation, we identify key areas for improvement, laying the foundation for future advancements.
FAQ
Common questions about the PhysicsFinals benchmark and leaderboard.