The authors evaluated stateful LLM agents in two Retail experiments and an Airline pilot using τ²-bench.
Task-specific plans and terminal checks create different value for LLM agents
Experiments in τ²-bench separate the benefits of meaningful planning guidance from those of independently verifying an agent’s final state.
Top University
Yukun Zhang · Kemu Xu · Yishen Chen
The Chinese University of Hong Kong · University of Edinburgh · The Chinese University of Hong Kong, Shenzhen
Research Digest··2 min read
Thread:Agent Harness Optimization
Zhang, Xu, and Chen compare task-specific plans with word-count-matched but shuffled policy text, finding that meaningful guidance raises oracle-verified success, especially on complex tasks.
Why this paper
From The Chinese University of Hong Kong and 2 others · Part of Agent Harness Optimization, now 61 papers
In one line
Prewritten plans improve oracle-verified success by 7.17 percentage points; a terminal verifier rejects 61% of invalid episodes at low cost.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§