The authors treat agent evaluation as selective prediction: a judge automatically labels only trajectories for which it is sufficiently confident, leaving the rest to humans.
Task-aware certificates safely automate part of LLM agent evaluation
Resampling at the task level let the authors certify automatic decisions despite correlations among agent trajectories.
Research Lab
Chengguang Gan · Yunhao Liang · Qinghao Zhang · Shiwen Ni
Independent Researcher · University of Chinese Academy of Sciences · Pusan National University · Shenzhen University of Advanced Technology
Research Digest··3 min read
Gan et al.
Why this paper
From University of Chinese Academy of Sciences and 3 others
In one line
A task-level bootstrap certificate ensures LLM judge error stays under budget, enabling safe selective automation of agent evaluation.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 4B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (6 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§