NAQD-Env represents plans using explicit dependencies among evidence, permissions, constraints, and actions.
Language agents rarely withdraw selectively when instructions or permissions change
In a synthetic benchmark, three open-weight models largely failed to suspend affected actions, preserve valid work, and resume after repair.
Abouzahra introduces NAQD-Env, a benchmark for testing whether language agents revise plans selectively when evidence, authorization, or constraints change.
Why this paper
Independent · Part of Agent Rule Compliance, now 23 papers
In one line
Language agents exhibit very low withdrawal recall and no valid resumption in the NAQD-Env benchmark.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§