← All threads

Open-Ended RL Reward Design

Methods for constructing rewards for reinforcement learning on tasks with no single correct answer, using relative comparisons or rubrics.

3 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.