FEAgent represents programs as typed, attributed program graphs capturing call-flow, control-flow, data-flow, type, import, and effect relations.
Program-graph evidence and blinded LLM surrogates detect code equivalence gaps
FEAgent reveals behavioral divergences missed by unit tests in 18% of benchmark pairs and 28% of test-passing patches.
Big Tech
Amit Kachroo · Like Hui · Haitao Mao · Yuhao Zhang · Nguyen Vo
AWS AI Labs
Research Digest··3 min read
The authors present FEAgent, a system that assesses functional equivalence of code by combining typed program-graph analysis with differential surrogate execution.
Why this paper
From AWS AI Labs
In one line
Graph-grounded, blinded surrogate execution finds behavioral differences that benchmark labels and passing unit tests miss while abstaining when evidence is incomplete.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (4 noted)
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§