The authors used the K framework to define executable semantics for Python and EVM.
Expert verification failures become reusable guidance for coding agents
A human-driven campaign using a K-framework semantics and verification kit achieved a 164/164 success rate on HumanEval and detected all defects in adversarial test packages
Academic
Yuqing Zhai · Xiaohong Chen · Lingming Zhang · Sriram Vishwanath · Grigore Rosu
University of Illinois Urbana-Champaign
Research Digest··3 min read
The authors developed a verification kit that captures expert diagnosis of verification failures into reusable procedures, combined with executable language definitions in the K framework.
Why this paper
From University of Illinois Urbana-Champaign
In one line
Expert diagnosis of verification failures can be captured as reusable guidance for coding agents.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§