The authors formulate the model-tool boundary as a finite robust control interface.
Agent certification can fail when proposal and history correlations are erased
The authors characterize exactly when state-only abstractions preserve safety certificates, and when observing a model’s proposal restores feasible intervention.
Academic
Hailin Zhong · Shengxin Zhu
Hong Kong Baptist University · Beijing Normal University
Research Digest··2 min read
Zhong and Zhu model an agent harness that observes a foundation model’s proposed action before deciding whether and how to execute it.
Why this paper
From Hong Kong Baptist University and Beijing Normal University
In one line
State-only abstraction of foundation-model agents preserves proposal coverage but can destroy certifiability unless every proposal-conditioned history fiber has a common robust-safe intervention.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§