The authors construct Dense Object Presence and Absence (DOPA), a dataset that exhaustively labels the deterministic presence or absence of concepts from an expanded object vocabulary.
Fine-Grained Rewards Reduce Vision-Language Hallucinations Without Silencing Valid Detail
The authors combine dense object annotations with subsentence-level reinforcement learning to reward accurate detail while penalizing localized hallucinations.
Research Lab
Xingming Long · Jie Zhang · Yuecong Min · Shiguang Shan · Xilin Chen
State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Zhongguancun Academy
Research Digest··3 min read
Thread:Process Agent Benchmarks
Long and colleagues address a failure mode of on-policy training for vision-language models: a model can reduce object hallucinations simply by making fewer claims, including fewer correct ones.
Why this paper
From Institute of Computing Technology, Chinese Academy of Sciences and 3 others · Part of Process Agent Benchmarks, now 9 papers
In one line
Fine-grained alignment using dense reward signals and subsentence-level credit assignment reduces object hallucination in vision-language models.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§