Benchmark of 82 unresolved problems in theoretical sciences tests AI beyond known results

GPT-6-Astra achieves a 14% mean solve rate, far above other models, with evaluator models judging correctness without reference answers

Research Lab
Zhiyi Li · Sihan Hu · Tianning Xiao · Xiansheng Cai · Xiaojun Tan · Youjin Deng · +1 more

University of Science and Technology of China · Chinese Academy of Sciences · Hefei National Laboratory · Hefei National Research Center for Physical Sciences at the Microscale · Institute for Advanced Algorithms Research

Research Digest··3 min read
The authors introduce OpenProblemBench, a benchmark of 82 open problems drawn from the mathematics and theoretical physics literature.

The authors curated 82 unresolved problems from the mathematics and theoretical physics literature, each including the problem statement, context, assumptions, known partial results, and conditions a resolution must satisfy.

Why this paper

From Chinese Academy of Sciences and 5 others

In one line

GPT-6-Astra achieves a 14% judged solve rate on open problems in mathematics and theoretical physics, outperforming other AI models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe