The authors repeatedly trained reasoning models using outcome-only rewards, including GRPO and stabilized PPO.
Reinforcement learning makes models’ admissions of failure hard to reproduce
Repeated training runs achieved similar task accuracy but produced sharply different rates of acknowledging unsuccessful solutions.
Big Tech
Steven Y. Feng · Noah D. Goodman · Michael C. Frank · Evan Hubinger · Paul C. Bogdan · Andrew Lampinen
Stanford University · Anthropic
Research Digest··2 min read
Feng et al.
Why this paper
From Anthropic and Stanford University
In one line
Failure disclosure in language models varies widely across reinforcement learning training runs, even when task accuracy remains stable.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§