The authors introduce TGRL, a method that for each prompt partitions the rollout group into high- and low-temperature subsets.
Temperature grouping boosts exploration efficiency in LLM reinforcement learning
By contrasting rollouts at different temperatures, TGRL allocates an explicit exploration signal without increasing sample budgets.
Research Lab
Zihan Lin · Xiaohan Wang · Jie Cao · Jiajun Chai · Wei Lin · Guojun Yin · +1 more
University of Chinese Academy of Sciences · Meituan · MAIS&NLPR, Institute of Automation, Chinese Academy of Sciences
Research Digest··2 min read
The authors propose Temperature-Grouped Reinforcement Learning (TGRL), which uses the reward contrast between high- and low-temperature rollouts to generate a per-token exploration signal, assigned via Jensen–Shannon divergence.
Why this paper
From University of Chinese Academy of Sciences and 2 others · Released code
In one line
Temperature-grouped RL converts temperature-driven reward differences into token-level credit, improving RLVR performance and reaching equivalent accuracy up to 36% faster without extra rollouts.
What it released
Code
What we could check
- ✓Code link in the paper (github.com)
- ·No weights link found
- ·No dataset link found
- ✓Compute or model size stated (params 32B)
- ·No stated limitations found
- ✓Reports numbers on named benchmarks (5 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§