How good are frontier models at physics?
4 hours ago
- Frontier language models' low scores on leading physics benchmarks may be misleading, as expert auditing reveals many incorrect evaluations stem from grader errors, faulty reference solutions, or ambiguous questions, not model reasoning flaws.
- Expert re-grading of six physics benchmarks (e.g., HLE-Physics, CMT-Benchmark, CritPt) substantially boosts model performance; for instance, GPT-5.6-Sol's mean@4 on HLE-Physics rises from 47.3% to 78.7%, and corrected pass@4 on CritPt reaches 94.4%.
- Current benchmarks significantly understate models' abilities on well-posed physics problems, indicating near-saturation on closed-ended tasks and emphasizing the need for more demanding, expert-validated evaluations.