The authors generated multiple candidate rollouts for each task, then used the same underlying model as both generator and verifier.
Agentic verification improves outputs on long-horizon workspace tasks
VeriHarness uses the generating model itself, environmental evidence and specialized verification routines to select and revise candidate outputs.
Big Tech
Caiqi Zhang · Rujun Han · Zifeng Wang · Zoey CuiZhu · Nigel Collier · Tomas Pfister · +1 more
Google Cloud AI Research · University of Cambridge
Research Digest··2 min read
Thread:Agent Self-Improvement
Zhang et al.
Why this paper
From Google Cloud AI Research and University of Cambridge · Part of Agent Self-Improvement, now 21 papers
In one line
VeriHarness uses the generator's own model as an agentic verifier to improve selection and revision of long-horizon task artifacts.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (3 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§