confidence calibration
- Papers
- 15
- Released code
- 2
- First seen
- Sept 2026
- Latest
- Oct 2026
15 papers in the last two months, against 0 in the two before.
The papers
Most central to this idea first, not most recent.
- Chinese Techcs.CL
Token-level certainty predicts LLM question difficulty better than response correctness
Shanghai Jiao Tong University, Alibaba Group · Oct 2026
- Research Labcs.AI
Jailbreak prompts improve uncertainty calibration in black-box reasoning models
Petscraft, Inria · Sept 2026
- Top Universitycs.AI
Teaching Models Confidence Can Shorten Reasoning Without Explicit Stop Training
University of Maryland, AI Foundations, Capital One · Sept 2026
- Chinese Techcs.LG
Differentiable confidence readouts improve reasoning accuracy and calibration together
University of Science and Technology of China, Qwen Business Unit of Alibaba · Sept 2026
- Research Labcs.CL
Companion model sharpens LLM confidence without harming task performance
State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences · Sept 2026
- Big Techcs.LG
Prediction error helps Decision Transformers reject unreliable rollout contexts
International Institute of Information Technology Bangalore, IBM Research Bangalore · Sept 2026
- Industrycs.LGcode
Training Forecasters to Search Improves Calibration and Reduces Tool Use
Future Principle · Oct 2026
- Research Labcs.LG
Event graphs improve LLM stock forecasts across two major markets
Zircon Security, Utrecht University · Oct 2026
- Top Universitycs.AI
Typed selective control cuts strong-model calls while preserving agent success
Nanyang Technological University · Sept 2026
- Academiccs.AI
Reasoning makes language models repeat the same wrong answers
Oklahoma State University · Sept 2026
- Research Labcs.AIcode
Agent performance depends more on system configuration than model size
Technical University of Munich, MCML · Oct 2026
- Big Techcs.CL
Reconstructing text under safety labels produces more reliable guard models
University of Neuchâtel, Delft University of Technology · Sept 2026
- Academiccs.CL
Short natural context can flip decision models to wrong answers
University of Southern California · Sept 2026
- Top Universitycs.CR
Unverified state text can flip calibrated model decisions
Institute of Information Engineering, Chinese Academy of Sciences, University of Chinese Academy of Sciences · Sept 2026
- Independentcs.LG
Fixed LLM judges can mismeasure upgrades as agents change
Sept 2026
Concepts are extracted from each paper and reused across the corpus, so this page grows on its own as the desk reads.