The authors fixed each task's failed tool observation and the evidence required for legitimate completion, separating reporting behavior from tool selection, retries, and environment dynamics.
Structured evidence reports reduce false success after tool failures
In a controlled benchmark, requiring agents to account for available evidence cut unsupported success claims while preserving useful responses.
Chinese Tech
Junru Zhu · Shiming Xie · Aime Lu Fan Chen · Xiaoqing Ding · Chunxin Tang · Ruoyu Qi · +1 more
Independent Researcher · Ant Group · Tsinghua University · University of Chicago · University of Waterloo
Research Digest··2 min read
Thread:Agent Rule Compliance
Zhu et al.
Why this paper
From Ant Group and 4 others · Part of Agent Rule Compliance, now 23 papers
In one line
Requiring structured evidence reports in tool-using language models lowers false-success rates from 22.8% to 0.8% while increasing useful responses.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§