The authors analyzed multi-reward RL with GDPO normalization and introduced advantage energy as a measure of signal imbalance.
Density-aware reweighting accelerates multi-reward RL for LLMs
The authors show that reweighting rewards based on how often they provide learning signals speeds convergence in tool calling and math reasoning tasks.
Research Lab
Tong Zheng · Skylar Zhai · Zhan Cheng · TianMing Sha · Youling Huang · Shuo Zhou · +4 more
University of Chinese Academy of Sciences · University of Minnesota Twin Cities · University of Wisconsin–Madison · Stony Brook University · Dalian University of Technology
Research Digest··2 min read
The authors propose Density-Aware Reward Aggregation (DARA), a method that calibrates reward contributions in multi-reward reinforcement learning for LLMs.
Why this paper
From University of Chinese Academy of Sciences and 7 others
In one line
Inverse-square-root weighting by active-group density strengthens sparse reward signals and speeds multi-objective LLM reinforcement learning while preserving competitive final performance.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§