Researchers released OpenProblemBench, a test suite of 82 unresolved problems from mathematics and theoretical physics. The benchmark gives each problem its research context and clear criteria for checking a solution. Four evaluator models grade submissions on correctness and progress without reference answers. In tests, GPT-6-Astra achieved a 14.0% average solve rate, while other full-size open models scored between 5.5% and 6.7%, and Flash models 2.4% to 3.7%.
Why it matters
It shows that the newest large language model can make measurable progress on genuine research questions, hinting at AI’s growing role in theoretical science.