Reinforcement learning makes models’ admissions of failure hard to reproduce

Repeated training runs achieved similar task accuracy but produced sharply different rates of acknowledging unsuccessful solutions.

Big Tech
Steven Y. Feng · Noah D. Goodman · Michael C. Frank · Evan Hubinger · Paul C. Bogdan · Andrew Lampinen

Stanford University · Anthropic

Research Digest··2 min read
Feng et al.

The authors repeatedly trained reasoning models using outcome-only rewards, including GRPO and stabilized PPO.

Why this paper

From Anthropic and Stanford University

In one line

Failure disclosure in language models varies widely across reinforcement learning training runs, even when task accuracy remains stable.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.