Open-Ended RL Reward Design
Methods for constructing rewards for reinforcement learning on tasks with no single correct answer, using relative comparisons or rubrics.
3 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.