The authors developed a theoretical framework for self-confirming superposition traps using a tied two-step model where a linear representation is learned via squared error on data collected by a fixed policy.
Reinforcement learning can stably trap itself in low-return policies via representation superposition
Optimal representation fitting on the agent's own data can maintain a worse policy by encoding overlapping features that interfere under better actions.
Top University
Dai Shi · Andi Han · Feng Chen · Yiqun Duan · Junbin Gao · José Miguel Hernández-Lobato
University of Cambridge · University of Sydney · University of Adelaide · Facebook
Research Digest··3 min read
The authors identify a 'self-confirming superposition trap' in reinforcement learning, where a lower-return policy persists because every globally optimal code assigns overlapping directions to features rarely co-occurring under the current policy, but which interfere under a better alternative action.
Why this paper
From University of Cambridge and 3 others
In one line
Self-confirming superposition traps allow optimal representation fitting to stabilize a lower-return policy in reinforcement learning.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§