Diverse supervised reasoning traces improve generalization after reinforcement learning

Selecting solutions with varied reasoning routes gave subsequent reinforcement learning a broader training signal and improved performance on unseen and harder problems.

Big Tech
Dylan Zhang · Mingyuan Wu · Jinning Li

University of Illinois Urbana-Champaign · Google

Research Digest··2 min read
Zhang, Wu and Li compare supervised fine-tuning datasets selected from the same candidate pools under matched training budgets, differing primarily in the diversity of their reasoning routes.

The authors developed a rule-based fingerprint that represents a solution's reasoning route, meaning its sequence of intermediate steps.

Why this paper

From Google and University of Illinois Urbana-Champaign

In one line

Selecting diverse reasoning traces for SFT improves post-RL generalization on puzzles and mathematics.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks (2 benchmarks)

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.