Yale University study reveals benchmark flaws in AI physics models
Summary
A new study involving Yale University and other top laboratories indicates that frontier models, such as GPT-5.6-Sol, are nearing the limits of current physics benchmark evaluations, which have often misrepresented their capabilities. Through expert audits of six popular physics benchmarks, researchers found that of 250 initially rejected cases, only 12 were legitimate model errors; the rest were due to ambiguous questions or incorrect grading, highlighting significant flaws in the benchmarking process. After corrections were made, GPT-5.6-Sol's performance score for the HLE-Physics benchmark rose from 47.3% to 78.7%. Despite these improvements in benchmark scoring, the study concludes that these models cannot consistently solve open theoretical physics problems, emphasizing the need for more rigorous evaluations that accurately reflect AI's physics capabilities.