TTCL derives two self-supervision signals from multiple model-generated answers: a correctness reward based on the majority answer and a calibration reward based on the empirical support of each answer.
Test-time calibration learning improves LLM accuracy and confidence without labels
A label-free framework adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data using self-supervision from multiple sampled responses.
Chinese Tech
Zizhuo Zhang · Xiong Peng · Jingwei Sun · Rong Yao · Shixiong Kai · Mingxuan Yuan · +1 more
Hong Kong Baptist University · Huawei Noah's Ark Lab
Research Digest··2 min read
The authors propose Test-Time Calibration Learning (TTCL), a method that jointly adapts an LLM's reasoning accuracy and confidence calibration to unseen target tasks without ground-truth labels.
Why this paper
From Huawei Noah's Ark Lab and Hong Kong Baptist University
In one line
An LLM's confidence can be calibrated at test time without labels using self-supervised rewards from multiple generated answers.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§