Tail-Safety Fine-Tuning
Methods for ensuring that fine-tuning large language models does not degrade safety performance beyond a specified threshold on worst-case tails, using chance-constrained optimization.
2 papers
Where this stands
The written synthesis of this thread is for subscribers. Subscribe.