Yang et al.
New benchmark tests automated theorem proving for research-level theoretical computer science
TCSAlgBench comprises 398 theorem-level challenges from recent STOC and COLT papers, and shows that current large language models achieve at most 25.4% coverage in verifier-accepted proofs.
Big Tech
Chutong Yang · Xiyuan Zhang · Yu Huang · Boran Han · Soonho Kong · Shuai Zhang · +4 more
University of Texas at Austin · Amazon · University of Pennsylvania
Research Digest··2 min read
The authors introduce TCSAlgBench, a benchmark and reusable pipeline for natural-language proof discovery in theoretical computer science.
Why this paper
From Amazon and 2 others
In one line
TCSAlgBench benchmarks automated proof discovery on 398 research-level TCS theorems; top models achieve about 25% coverage.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§