Token-aware latent refinement improves vision-language reasoning at inference time

The method directs visual and reasoning feedback to different hidden-state tokens while keeping model parameters frozen.

Chinese Tech
Hao-Xuan Ma · Yihao Liu · Yutao Sun · Yanting Miao · Mengyu Zhou · YiCheng Xiao · +5 more

Qwen Business Unit of Alibaba · Nanjing University · National Key Laboratory for Novel Software Technology, Nanjing University · Zhejiang University · University of Waterloo

Research Digest··2 min read
Ma et al.

The authors begin with an initial model-generated trajectory and optimize a short prefix of its hidden states without changing the model's parameters.

Why this paper

From Chinese Academy of Sciences and 7 others · Released code

In one line

Token-disentangled latent test-time scaling improves vision-language reasoning by routing visual and reasoning feedback to different token roles.

What it released

Code

What we could check

  • ✓Code link in the paper (github.com)
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ✓Reports numbers on named benchmarks

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.