Steering hidden states stabilizes vision-language reasoning against subtle visual changes

FlipDir estimates a low-rank activation subspace from contrastive pairs and selectively steers only uncertain decoding steps, outperforming baselines across 18 evaluation settings.

Big Tech
Yeonsung Jung · Joonhyun Jeong · Hoang Pham · Joowon Kim · Yoonsik Park · Viet Dac Lai · +1 more

KAIST · NAVER Cloud · The Ohio State University · Adobe Research · AITRICS

Research Digest··3 min read
The authors identify answer flips—cases where nearly identical images cause vision-language models to produce different answers—and propose FlipDir, a training-free inference-time method that estimates a low-rank subspace of flip-inducing activations and selectively steers hidden states during decoding.

, exposure shifts, JPEG compression, brightness changes) that leave images nearly identical to human observers but alter the model's final prediction.

Why this paper

From Adobe Research and 4 others

In one line

FlipDir is a training-free method that steers VLM hidden states to recover answers flipped by subtle visual variations.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.