The authors cast collaborative model training as a leader-follower game.
Adaptive instruction selection helps language models learn from peers
A bandit-guided training curriculum improved collaborative alignment by directing model competitions toward instructions that produced useful preference signals.
Big Tech
Christina Hahn · Shangbin Feng · Dean Light · Swastik Roy · Hila Gonen · Yulia Tsvetkov
University of Washington · Amazon · University of British Columbia
Research Digest··3 min read
Hahn et al.
Why this paper
From Amazon and 2 others
In one line
Adaptive instruction selection via a Stackelberg game improves multi-LLM collaborative alignment across benchmarks.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§