Post-training alignment makes language models more confidently wrong

Across five model families, instruction tuning sharply increased high-confidence factual errors, while constraining preference-optimization margins reduced them.

Research Lab
Qingjia Huang · Yakai Li · Jianguo Wu · Qihang Zhou · Aimin Yu · Xiaoqi Jia · +2 more

Institute of Information Engineering, Chinese Academy of Sciences · School of Cyber Security, University of Chinese Academy of Sciences

Research Digest··2 min read
Huang and colleagues compared base and aligned language models on long-tail factual questions, then traced how confidence in incorrect answers developed across model layers.

Why this paper

From Institute of Information Engineering, Chinese Academy of Sciences and School of Cyber Security, University of Chinese Academy of Sciences · Part of Safety Training Side Effects, now 5 papers

In one line

Alignment training, not just missing knowledge, is a primary cause of LLMs' confident hallucinations, multiplying high-confidence factual errors by 10x to 35x; bounded-margin DPO cuts them.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.