, exposure shifts, JPEG compression, brightness changes) that leave images nearly identical to human observers but alter the model's final prediction.
Steering hidden states stabilizes vision-language reasoning against subtle visual changes
FlipDir estimates a low-rank activation subspace from contrastive pairs and selectively steers only uncertain decoding steps, outperforming baselines across 18 evaluation settings.
Big Tech
Yeonsung Jung · Joonhyun Jeong · Hoang Pham · Joowon Kim · Yoonsik Park · Viet Dac Lai · +1 more
KAIST · NAVER Cloud · The Ohio State University · Adobe Research · AITRICS
Research Digest··3 min read
Thread:VLM Hallucination Circuits
The authors identify answer flips—cases where nearly identical images cause vision-language models to produce different answers—and propose FlipDir, a training-free inference-time method that estimates a low-rank subspace of flip-inducing activations and selectively steers hidden states during decoding.
Why this paper
From Adobe Research and 4 others
In one line
FlipDir is a training-free method that steers VLM hidden states to recover answers flipped by subtle visual variations.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§