The authors evaluated RL-trained Qwen and Gemma models across multiple reasoning domains, examining both Pass@1 and Pass@N.
Reinforcement learning reshapes which reasoning strategies language models prefer
Experiments and theory suggest RL can improve accuracy by reweighting existing strategies, while either broadening or narrowing the problems solved through repeated sampling.
Top University
Ziheng Cheng · Yixiao Huang · Hanlin Zhu · Somayeh Sojoudi
University of California, Berkeley
Research Digest··3 min read
Cheng and colleagues studied reinforcement learning across Qwen and Gemma models, comparing performance and sampled solution traces before and after training.
Why this paper
From University of California, Berkeley
In one line
Reinforcement learning primarily reweights existing reasoning strategies rather than expanding an LLM's reasoning boundary.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§