The authors introduce the Pairwise Faithfulness Benchmark, or PFaithBench.
Faithful LLMs Must Adapt Their Responses When Context Changes
PFaithBench tests whether models answer a question with supporting evidence but abstain when evidence for the same question is removed.
Chinese Tech
Zizhuo Zhang · Xiong Peng · Jingwei Sun · Rong Yao · Borui Jiang · Bo Han
Hong Kong Baptist University · Huawei Noah’s Ark Lab
Research Digest··2 min read
Zhang et al.
Why this paper
From Huawei Noah’s Ark Lab and Hong Kong Baptist University
In one line
LLMs are biased toward answering rather than abstaining when context is insufficient, and training data composition drives the answering-abstaining tradeoff.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§