Pluralistic Preference Optimization, or PlurPO, begins with prompts describing interpersonal conflicts.
Training models on multiple viewpoints reduces excessive user affirmation
Pluralistic Preference Optimization teaches language models to favor advice acceptable to every simulated stakeholder in an interpersonal conflict.
Big Tech
Stephane Hatgis-Kessell · Myra Cheng · Xiaoxuan Hou · Qian Hu · Rahul Gupta · Natasha Jaques · +1 more
Stanford University · Amazon · University of Washington
Research Digest··3 min read
Hatgis-Kessell et al.
Why this paper
From Amazon and 2 others
In one line
Pluralistic Preference Optimization reduces social sycophancy by training models to consider multiple stakeholder perspectives.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (3 noted)
- ✓Reports numbers on named benchmarks (4 benchmarks)
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§