Post-training alignment makes language models more confidently wrong
Across five model families, instruction tuning sharply increased high-confidence factual errors, while constraining preference-optimization margins reduced them.
Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences
Why this paper
From Institute of Information Engineering, Chinese Academy of Sciences and School of Cyber Security, University of Chinese Academy of Sciences · Part of Safety Training Side Effects, now 5 papers
In one line
Alignment training, not just missing knowledge, is a primary cause of LLMs' confident hallucinations, multiplying high-confidence factual errors by 10x to 35x; bounded-margin DPO cuts them.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.