The authors trained multi-agent systems on CodeContests and Polaris-Dataset-53K, using specialized coder and tester roles for programming and reasoner and tool-user roles for mathematics.
Teacher-guided training improves specialization and coordination in multi-agent models
MAS-OPD trains interacting small models with token-level teacher feedback that distinguishes role-specific behavior and attributes collaboration failures.
Chinese Tech
Qiyong Zhong · Mao Zheng · Mingyang Song · Houcheng Jiang · Jiajie Su · Huwei Ji · +2 more
University of Science and Technology of China · Tencent · Zhejiang University · National University of Singapore
Research Digest··2 min read
Zhong et al.
Why this paper
From Tencent and 3 others
In one line
MAS-OPD achieves the highest mean score on code and math benchmarks by using on-policy distillation with role-advantage specialization and privileged attribution.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§