The authors define preference-based formulations for both decentralized systems, where agents act from local context, and centralized systems, where communication or a manager coordinates the team.
Preference learning improves collaboration among teams of language-model agents
The proposed MAPL framework trains centralized and decentralized agent teams from comparative feedback instead of relying on handcrafted task rewards.
Independent
Shuo Liu · Xinzichen Li · Tianle Chen · Christopher Amato
Research Digest··2 min read
Thread:Multi-Agent Coordination
Liu et al.
Why this paper
Independent · Part of Multi-Agent Coordination, now 34 papers
In one line
Multi-agent preference learning (MAPL) improves LLM collaboration quality and efficiency while approaching performance of MARL with fixed rewards.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§