Co-evolving agent harnesses improve abstention without sacrificing task completion

HERA iteratively adapts an agent’s control system and its training environments, producing a harness that better distinguishes solvable tasks from those requiring abstention.

Top University
Han Luo · Bingbing Wen · Guang Yang · Zora Zhiruo Wang · Pan Lu · Lucy Lu Wang

University of Washington · University of Leeds · Carnegie Mellon University · Stanford University · Allen Institute for AI

Research Digest··3 min read
Luo and colleagues introduce HERA, a framework that generates matched feasible and infeasible tool-use tasks, then uses observed failures to alternately improve the agent harness and create new challenges.

The authors first constructed executable tasks with verified reference solutions.

Why this paper

From Allen Institute for AI and 4 others

In one line

Co-evolving an agent harness with adversarially mutated environments improves abstention, task completion, cross-model transfer, and cost efficiency in tool-using LLM agents.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.