The authors formalized a family of algorithms called RE(S), where S ≥ 1 is the number of gradient steps between updates of the rollout distribution.
Off-policy sampling avoids suboptimal traps in reward-guided self-training
A theoretical analysis of RE(S) shows that updating the rollout distribution every S gradient steps can escape local detours and achieve faster global convergence than on-policy REINFORCE when starting from a weak policy.
Chinese Tech
Zhiwei Wang · Yanxi Chen · Yaliang Li · Bolin Ding
Tsinghua University · Alibaba Group
Research Digest··3 min read
The authors study RE(S), a generalization of REINFORCE that updates the rollout distribution once every S gradient steps.
Why this paper
From Alibaba Group and Tsinghua University
In one line
For softmax multi-arm bandits, infrequently refreshed rollout data still converges at Θ(1/T) and can outperform on-policy REINFORCE from weak initial policies.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§