Reinforcement learning reshapes which reasoning strategies language models prefer

Experiments and theory suggest RL can improve accuracy by reweighting existing strategies, while either broadening or narrowing the problems solved through repeated sampling.

Top University
Ziheng Cheng · Yixiao Huang · Hanlin Zhu · Somayeh Sojoudi

University of California, Berkeley

Research Digest··3 min read
Cheng and colleagues studied reinforcement learning across Qwen and Gemma models, comparing performance and sampled solution traces before and after training.

The authors evaluated RL-trained Qwen and Gemma models across multiple reasoning domains, examining both Pass@1 and Pass@N.

Why this paper

From University of California, Berkeley

In one line

Reinforcement learning primarily reweights existing reasoning strategies rather than expanding an LLM's reasoning boundary.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.