# OpenProblemBench benchmark released for AI on unsolved math and physics questions

> The new OpenProblemBench set contains 82 open problems; GPT-6-Astra solved 14% of them, far above other models.

Oossa · 2026-10-09 · https://oossa.com/en/openproblembench-benchmark-released-for-ai-on-unsolved-math-and-physics-question

Researchers released OpenProblemBench, a test suite of 82 unresolved problems from mathematics and theoretical physics. The benchmark gives each problem its research context and clear criteria for checking a solution. Four evaluator models grade submissions on correctness and progress without reference answers. In tests, GPT-6-Astra achieved a 14.0% average solve rate, while other full-size open models scored between 5.5% and 6.7%, and Flash models 2.4% to 3.7%.

## The facts

- 82 open problems in the benchmark
- GPT-6-Astra mean solve rate: 14.0%

## Why it matters

It shows that the newest large language model can make measurable progress on genuine research questions, hinting at AI’s growing role in theoretical science.

## Sources & references

1. [OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences](https://arxiv.org/abs/2610.11118) – arXiv, 2026-10-09

Last updated: 2026-10-09
