Reinforcement learning with verifiable reward causes unbounded language drift in LLM reasoning chains

Sullivan and Koller prove theoretically that RLVR training permits unlimited linguistic deviation, while supervised fine-tuning bounds it, and show that constraining drift necessarily constrains expected reward.

Academic
Michael Sullivan · Alexander Koller

Saarland University

Research Digest··2 min read
The authors investigate language drift in LLM chains of thought during reinforcement learning with verifiable reward (RLVR) post-training.

Sullivan and Koller first provide a theoretical analysis of language drift under RLVR and supervised fine-tuning (SFT).

Why this paper

From Saarland University

In one line

Reinforcement learning with verifiable reward causes unbounded language drift that cannot be prevented without harming performance.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.