The authors created RSI-Master around two components.
Structured experiment orchestration helps agents improve models without benchmark hacking
RSI-Master combines constrained experiment tools with reviewer-guided branching to explore post-training strategies while preserving traceable records.
Top University
Yaxin Du · Xiyuan Yang · Zhifan Zhou · Yujie Ge · Cheng Wang · Jiajun Wang · +7 more
Shanghai Jiao Tong University · Carnegie Mellon University · University of Waterloo
Research Digest··2 min read
Thread:Agent Self-Improvement
Du and colleagues built an agent system that autonomously develops post-training strategies for language models using a structured, branching research process.
Why this paper
From Shanghai Jiao Tong University and 2 others · Part of Agent Self-Improvement, now 21 papers
In one line
RSI-Master uses an Experiment OS and reviewer-guided research DAG to stop hacking and strategy lock-in, improving autonomous post-training model development.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 4B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§