Sycophancy Mitigation
Training methods and protocols that reduce language models' tendency to affirm user positions even when harmful.
2 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.
News without the noise
Training methods and protocols that reduce language models' tendency to affirm user positions even when harmful.
2 papers
The written synthesis of this thread is for subscribers. Subscribe.