The authors distinguish internal probabilities, derived from a model’s sampling distribution or logits, from verbalized probabilities, obtained by asking the model to state its confidence.
Stated confidence tracks language models’ internal answer probabilities
Controlled interventions show that frequency evidence and explicit probability claims affect both forms of uncertainty readout.
Big Tech
Sinead Williamson · Jiaxuan Li · Nick Foti · Russ Webb · Masha Fedzechkina
Apple
Research Digest··3 min read
Williamson and colleagues test whether a model’s stated confidence reflects the probabilities encoded in its next-token distribution.
Why this paper
From Apple
In one line
Verbalized probabilities can be used to probe a large language model's internal probability distribution.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§