The authors propose PoEM, a framework that takes a set of RL post-trained policies (each optimized for a different reward function) and a new target reward, and outputs an approximation of the policy that would result from running RL on that new reward.
A framework predicts RL post-training outcomes without running reinforcement learning
PoEM composes existing post-trained policies to approximate the policy for a new reward function, eliminating the need for additional RL training.
Top University
Kimia Hamidieh · Giannis Daras · Antonio Torralba
MIT CSAIL
Research Digest··3 min read
Thread:RL for Tool Agents
The authors introduce PoEM, a method that predicts the output of reinforcement learning (RL) on a new reward function by linearly combining log-policies from models already post-trained on other rewards.
Why this paper
From MIT CSAIL · Part of RL for Tool Agents, now 3 papers
In one line
PoEM predicts reinforcement learning outcomes for a new reward using existing policies without additional RL training.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§