The authors address two limitations in multimodal speculative decoding.
HAWK boosts speculative decoding for LVLMs via target hidden states and prediction-shift training
The method uses representation similarity to select target layers, compresses visual hidden states for the drafter, and trains it to match target logit shifts after its own proposals.
Big Tech
Wenhan Yang · Anirudh Rao · Ashwin Chandra
University of California, Los Angeles · Samsung Research America
Research Digest··3 min read
The authors propose HAWK, a speculative decoding framework for large vision-language models.
Why this paper
From Samsung Research America and University of California, Los Angeles
In one line
HAWK improves multimodal speculative decoding by feeding drafters compressed target visual hidden states and supervising prediction shifts, raising acceptance length and speedup.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ✓Reports numbers on named benchmarks
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§