Aligning policy and reward manifolds stabilizes LLM reasoning in RL

DCRL combines syllogistic prompt evolution with joint reward-policy updates, beating both rule-based and learned reward baselines on reasoning benchmarks.

Chinese Tech
Henan Sun · Zehua Li · Haitao Hu · Qifan Zhang · Jianfeng Zhang · Nuo Chen · +1 more

The Hong Kong University of Science and Technology (Guangzhou) · Huawei Noah's Ark Lab · Tencent HY

Research Digest··2 min read
The authors propose Decoupling and Coupling Reinforcement Learning (DCRL), which reframes LLM reasoning as a coupled geometric manifold and aligns policy and reward models during training.

The authors model LLM reasoning as three interdependent geometric sub-manifolds: logical deduction, evaluation, and representation.

Why this paper

From Huawei Noah's Ark Lab and 2 others

In one line

DCRL aligns policy and reward manifolds to improve LLM reasoning, enabling a 4B model to surpass a 32B baseline.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.