← All threads

Sycophancy Mitigation

Training methods and protocols that reduce language models' tendency to affirm user positions even when harmful.

2 papers

Where this stands

The written synthesis of this thread is for subscribers. Subscribe.