The authors curated 82 unresolved problems from the mathematics and theoretical physics literature, each including the problem statement, context, assumptions, known partial results, and conditions a resolution must satisfy.
Benchmark of 82 unresolved problems in theoretical sciences tests AI beyond known results
GPT-6-Astra achieves a 14% mean solve rate, far above other models, with evaluator models judging correctness without reference answers
Research Lab
Zhiyi Li · Sihan Hu · Tianning Xiao · Xiansheng Cai · Xiaojun Tan · Youjin Deng · +1 more
University of Science and Technology of China · Chinese Academy of Sciences · Hefei National Laboratory · Hefei National Research Center for Physical Sciences at the Microscale · Institute for Advanced Algorithms Research
Research Digest··3 min read
The authors introduce OpenProblemBench, a benchmark of 82 open problems drawn from the mathematics and theoretical physics literature.
Why this paper
From Chinese Academy of Sciences and 5 others
In one line
GPT-6-Astra achieves a 14% judged solve rate on open problems in mathematics and theoretical physics, outperforming other AI models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§