Entesari and Fazlyab frame safety-preserving fine-tuning as a reliability-constrained optimization problem.
Limiting tail safety failures during fine-tuning beats average-loss defenses
The authors propose a chance-constrained optimization method that caps the fraction of safety examples whose degradation exceeds a threshold, and show it outperforms existing defenses on harmful fine-tuning tasks.
Top University
Taha Entesari · Mahyar Fazlyab
Johns Hopkins University
Research Digest··3 min read
Thread:Tail-Safety Fine-Tuning
The authors introduce a chance-constrained formulation for fine-tuning large language models that bounds the fraction of safety examples whose performance degrades beyond a specified threshold, rather than minimizing average safety loss.
Why this paper
From Johns Hopkins University
In one line
Chance-constrained fine-tuning limits the fraction of safety examples that degrade beyond a threshold, outperforming average-risk methods.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§