The authors tested three student models on AppWorld, a benchmark for agents operating applications, and SWE-bench Verified, which evaluates software issue resolution.
Privileged self-practice beats self-distillation for multi-turn language-model agents
A gated training method uses privileged instructions to generate useful practice trajectories without forcing agents to imitate an information-rich teacher.
Big Tech
Texas A&M University · AWS AI, Amazon
Research Digest··2 min read
Su et al.
Why this paper
From AWS AI, Amazon and Texas A&M University
In one line
Privileged Self-Practice injects task-specific instructions into failing agent rollouts to improve multi-turn agent training.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§