The authors model LLM reasoning as three interdependent geometric sub-manifolds: logical deduction, evaluation, and representation.
Aligning policy and reward manifolds stabilizes LLM reasoning in RL
DCRL combines syllogistic prompt evolution with joint reward-policy updates, beating both rule-based and learned reward baselines on reasoning benchmarks.
Chinese Tech
Henan Sun · Zehua Li · Haitao Hu · Qifan Zhang · Jianfeng Zhang · Nuo Chen · +1 more
The Hong Kong University of Science and Technology (Guangzhou) · Huawei Noah's Ark Lab · Tencent HY
Research Digest··2 min read
The authors propose Decoupling and Coupling Reinforcement Learning (DCRL), which reframes LLM reasoning as a coupled geometric manifold and aligns policy and reward models during training.
Why this paper
From Huawei Noah's Ark Lab and 2 others
In one line
DCRL aligns policy and reward manifolds to improve LLM reasoning, enabling a 4B model to surpass a 32B baseline.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§