Teacher-guided training improves specialization and coordination in multi-agent models

MAS-OPD trains interacting small models with token-level teacher feedback that distinguishes role-specific behavior and attributes collaboration failures.

Chinese Tech
Qiyong Zhong · Mao Zheng · Mingyang Song · Houcheng Jiang · Jiajie Su · Huwei Ji · +2 more

University of Science and Technology of China · Tencent · Zhejiang University · National University of Singapore

Research Digest··2 min read
Zhong et al.

The authors trained multi-agent systems on CodeContests and Polaris-Dataset-53K, using specialized coder and tester roles for programming and reasoner and tool-user roles for mathematics.

Why this paper

From Tencent and 3 others

In one line

MAS-OPD achieves the highest mean score on code and math benchmarks by using on-policy distillation with role-advantage specialization and privileged attribution.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.