Stated confidence tracks language models’ internal answer probabilities

Controlled interventions show that frequency evidence and explicit probability claims affect both forms of uncertainty readout.

Big Tech
Sinead Williamson · Jiaxuan Li · Nick Foti · Russ Webb · Masha Fedzechkina

Apple

Research Digest··3 min read
Williamson and colleagues test whether a model’s stated confidence reflects the probabilities encoded in its next-token distribution.

The authors distinguish internal probabilities, derived from a model’s sampling distribution or logits, from verbalized probabilities, obtained by asking the model to state its confidence.

Why this paper

From Apple

In one line

Verbalized probabilities can be used to probe a large language model's internal probability distribution.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.