Static-analysis checkers encode defect patterns as tool-specific rules, but manually designing them requires substantial expert effort.
Benchmark shows coding agents struggle to synthesize static-analysis checkers
CheckerBench spans 300 CVE-derived tasks across five language ecosystems; the best agent Pass@1 is 45.33%.
Top University
Hang He · Li Wang · Hao Chen · Yuchen Shao · Yuling Shi · Lisheng Wang · +6 more
East China Normal University · Humanlaya Data · Shanghai Jiao Tong University · Peking University · Shanghai Innovation Institute
Research Digest··2 min read
The authors introduce CheckerBench, an executable benchmark of 300 real-world tasks for long-horizon static-analysis checker synthesis, and CheckerLab, a unified evaluation framework.
Why this paper
From Shanghai Jiao Tong University and 4 others
In one line
Current coding agents reliably synthesize static-analysis checkers on fewer than half of CheckerBench tasks, with the best configuration achieving 45.33% Pass@1.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§