Reinforcement learning can stably trap itself in low-return policies via representation superposition

Optimal representation fitting on the agent's own data can maintain a worse policy by encoding overlapping features that interfere under better actions.

Top University
Dai Shi · Andi Han · Feng Chen · Yiqun Duan · Junbin Gao · José Miguel Hernández-Lobato

University of Cambridge · University of Sydney · University of Adelaide · Facebook

Research Digest··3 min read
The authors identify a 'self-confirming superposition trap' in reinforcement learning, where a lower-return policy persists because every globally optimal code assigns overlapping directions to features rarely co-occurring under the current policy, but which interfere under a better alternative action.

The authors developed a theoretical framework for self-confirming superposition traps using a tied two-step model where a linear representation is learned via squared error on data collected by a fixed policy.

Why this paper

From University of Cambridge and 3 others

In one line

Self-confirming superposition traps allow optimal representation fitting to stabilize a lower-return policy in reinforcement learning.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.