A framework predicts RL post-training outcomes without running reinforcement learning

PoEM composes existing post-trained policies to approximate the policy for a new reward function, eliminating the need for additional RL training.

Top University
Kimia Hamidieh · Giannis Daras · Antonio Torralba

MIT CSAIL

Research Digest··3 min read
The authors introduce PoEM, a method that predicts the output of reinforcement learning (RL) on a new reward function by linearly combining log-policies from models already post-trained on other rewards.

The authors propose PoEM, a framework that takes a set of RL post-trained policies (each optimized for a different reward function) and a new target reward, and outputs an approximation of the policy that would result from running RL on that new reward.

Why this paper

From MIT CSAIL · Part of RL for Tool Agents, now 3 papers

In one line

PoEM predicts reinforcement learning outcomes for a new reward using existing policies without additional RL training.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.