TIPS prompts a language model to generate a chain of thought, step-level validity labels and a final outcome label.
Outcome-only reinforcement learning can teach models to judge reasoning steps
TIPS trained generative process reward models using final-outcome labels alone, yet improved their ability to identify faulty intermediate steps.
Top University
Shengda Fan · Xin Cong · Zhong Zhang · Haotian Chen · Yankai Lin
Renmin University of China · Tsinghua University · University of Electronic Science and Technology of China · Shanghai Jiao Tong University
Research Digest··2 min read
Fan et al.
Why this paper
From Tsinghua University and 3 others
In one line
Outcome-only RL can induce step-level process supervision in generative reward models, reaching 85.2 F1 on ProcessBench with 3.2K outcome-labeled trajectories.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§