Fine-Grained Rewards Reduce Vision-Language Hallucinations Without Silencing Valid Detail

The authors combine dense object annotations with subsentence-level reinforcement learning to reward accurate detail while penalizing localized hallucinations.

Research Lab
Xingming Long · Jie Zhang · Yuecong Min · Shiguang Shan · Xilin Chen

State Key Laboratory of AI Safety · Institute of Computing Technology, Chinese Academy of Sciences · University of Chinese Academy of Sciences · Zhongguancun Academy

Research Digest··3 min read
Long and colleagues address a failure mode of on-policy training for vision-language models: a model can reduce object hallucinations simply by making fewer claims, including fewer correct ones.

The authors construct Dense Object Presence and Absence (DOPA), a dataset that exhaustively labels the deterministic presence or absence of concepts from an expanded object vocabulary.

Why this paper

From Institute of Computing Technology, Chinese Academy of Sciences and 3 others · Part of Process Agent Benchmarks, now 9 papers

In one line

Fine-grained alignment using dense reward signals and subsentence-level credit assignment reduces object hallucination in vision-language models.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.