The authors replaced text-generated numerical confidence with a deterministic readout from the model's output distribution.
Differentiable confidence readouts improve reasoning accuracy and calibration together
CREDO derives confidence from reserved-token probabilities and trains it directly, avoiding the noise and discreteness of confidence sampled as text.
Chinese Tech
Chenxiao Fan · Chongming Gao · Gangyi Zhang · Leyang Shen · Yaxin Gong · Jiamin Wang · +5 more
University of Science and Technology of China · Qwen Business Unit of Alibaba · National University of Singapore
Research Digest··2 min read
Fan and colleagues introduce CREDO, a method for jointly improving correctness and confidence calibration in reasoning models trained with reinforcement learning from verifiable rewards.
Why this paper
From Qwen Business Unit of Alibaba and 2 others
In one line
CREDO improves reasoning accuracy and confidence calibration by replacing sampled confidence text with a deterministic, differentiable token-probability readout.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§