New benchmark tests automated theorem proving for research-level theoretical computer science

TCSAlgBench comprises 398 theorem-level challenges from recent STOC and COLT papers, and shows that current large language models achieve at most 25.4% coverage in verifier-accepted proofs.

Big Tech
Chutong Yang · Xiyuan Zhang · Yu Huang · Boran Han · Soonho Kong · Shuai Zhang · +4 more

University of Texas at Austin · Amazon · University of Pennsylvania

Research Digest··2 min read
The authors introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery in theoretical computer science.

Yang et al.

Why this paper

From Amazon and 2 others

In one line

TCSAlgBench benchmarks automated proof discovery on 398 research-level TCS theorems; top models achieve about 25% coverage.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.