The authors developed a rule-based fingerprint that represents a solution's reasoning route, meaning its sequence of intermediate steps.
Diverse supervised reasoning traces improve generalization after reinforcement learning
Selecting solutions with varied reasoning routes gave subsequent reinforcement learning a broader training signal and improved performance on unseen and harder problems.
Big Tech
Dylan Zhang · Mingyuan Wu · Jinning Li
University of Illinois Urbana-Champaign · Google
Research Digest··2 min read
Zhang, Wu and Li compare supervised fine-tuning datasets selected from the same candidate pools under matched training budgets, differing primarily in the diversity of their reasoning routes.
Why this paper
From Google and University of Illinois Urbana-Champaign
In one line
Selecting diverse reasoning traces for SFT improves post-RL generalization on puzzles and mathematics.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (2 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§