The authors investigated what LLMs learn when trained to estimate their own performance (metacognitive monitoring).
LLMs trained to predict their own accuracy learn two distinct types of confidence
Fine-tuned confidence tracks true accuracy on familiar questions but shifts to output consistency on novel topics.
Research Lab
Nicolas Yax · Stefano Palminteri · Pierre-Yves Oudeyer
INSERM · ENS PSL · Inria
Research Digest··2 min read
The authors fine-tuned 10 open-weight LLMs to predict their correctness on factual multiple-choice questions before answering.
Why this paper
From Inria and 2 others
In one line
LLMs trained to predict their own accuracy learn to track output consistency, not true accuracy, on unfamiliar questions.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§