Test-time calibration learning improves LLM accuracy and confidence without labels

A label-free framework adapts reasoning accuracy and verbalized confidence directly on unlabeled target-task data using self-supervision from multiple sampled responses.

Chinese Tech
Zizhuo Zhang · Xiong Peng · Jingwei Sun · Rong Yao · Shixiong Kai · Mingxuan Yuan · +1 more

Hong Kong Baptist University · Huawei Noah's Ark Lab

Research Digest··2 min read
The authors propose Test-Time Calibration Learning (TTCL), a method that jointly adapts an LLM's reasoning accuracy and confidence calibration to unseen target tasks without ground-truth labels.

TTCL derives two self-supervision signals from multiple model-generated answers: a correctness reward based on the majority answer and a calibration reward based on the empirical support of each answer.

Why this paper

From Huawei Noah's Ark Lab and Hong Kong Baptist University

In one line

An LLM's confidence can be calibrated at test time without labels using self-supervised rewards from multiple generated answers.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.