The authors represent every reasoning prefix with a binary latent state: valid, meaning the current derivation contains no unresolved error, or invalid.
Propagating reasoning states lets outcome labels improve process rewards
The proposed model tracks when reasoning becomes invalid or recovers, allowing final-answer labels to help train step-level evaluations.
Academic
Kai Gan · Zi-Hao Zhou · Bo Ye · Jian Zhao · Min-Ling Zhang · Tong Wei
Southeast University · Key Laboratory of Computer Network and Information Integration (Southeast University) · Ministry of Education, China · Zhongguancun Academy · Zhongguancun Institute of Artificial Intelligence
Research Digest··3 min read
Gan and colleagues introduce Reasoning State Propagation (RSP), a process reward model that explicitly connects the validity of successive steps in a reasoning trajectory.
Why this paper
From Ministry of Education, China and 4 others
In one line
Reasoning State Propagation propagates binary validity states through break and repair probabilities, combining process and outcome supervision.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§