The authors extracted 509 task directions from SkillsBench using only information visible to the agent: prompts, workspace contents, and injected skill specifications.
External specifications, not agents, should authorize task completion claims
SpecHarness converts agent-visible instructions into tracked obligations and requires admissible evidence before verifiable requirements can be marked complete.
Independent
Haiqing Li · Xin Ma · Yinhao Wu · Wenliang Zhong · Feng Jiang · Thao M. Dang · +4 more
Research Digest··2 min read
Thread:Agent Rule Compliance
Li and colleagues examine a structural weakness in language-model agents: the same model often performs a task, evaluates its own work, and declares completion.
Why this paper
Independent · Part of Agent Rule Compliance, now 17 papers
In one line
Large language model agents should not declare their own completion: specification-governed state must be established by admissible evidence from qualified providers.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§