Agent distillation improves when supervision follows outcomes, not model disagreement

Outcome-guided weighting identifies which teacher interventions actually help the current student complete multi-turn tasks.

Chinese Tech
Tong Zhang · Zhou Liu · Yihao Liu · Jiahua Bao · Xuchen Li · Honglin Lin · +5 more

Qwen Large Model Application Team, Alibaba · Peking University · University of Chinese Academy of Sciences · Shanghai Jiao Tong University

Research Digest··2 min read
Zhang et al.

The authors analyzed teacher interventions in ScienceWorld and WebShop by branching from the same interaction state.

Why this paper

From University of Chinese Academy of Sciences and 3 others

In one line

Outcome-calibrated teacher supervision improves multi-turn agent distillation because token-level teacher-student gaps alone do not predict which interventions improve task success.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.