72 papers this week in Vision12 active threadsbusiest: Context Engineering for Agentsdaily arXiv scan · 6am Brisbane

Vision research

Latest Paper· Generative media

Adaptive reference attention reuse speeds up conditioned diffusion transformers

The authors propose RefAdapt-DiT, a training-free framework that adaptively controls joint attention computation between reference images and target generation in diffusion transformers. By monitoring reference drift and target-to-reference attention, it selectively reuses stale reference states, achieving speedups of up to 3.54× on image editing with minimal quality loss.

Jian Tang, Jiawei Fan, Qiannan Zhou +3
Vision

Decoupled temporal axis yields motion-prioritized video representations efficiently

The authors conduct a controlled architecture-objective study and propose TT-VidT, which uses a DINOv3-initialized ViT-B/16 per-frame spatial path and a compact Temporal Transfer Layer, trained with Diff Compression to reconstruct frames from a first-frame anchor and motion tokens. They show TT-VidT leads fine-tuning on Jester, Something-Something V2, ARID, and Diving48, improving 54–121% over non-TT methods, with 48–55% fewer encoder FLOPs than DisMo and VideoMAE/V-JEPA2.

today
Vision

Pruning visual tokens and decoder sublayers cuts multimodal model computation

Chen and colleagues target two sources of multimodal model overhead: redundant visual tokens and unnecessary computation inside the language-model decoder. Their training-free SPIDER framework retains complementary visual information from middle and deep vision layers, then selectively skips attention or feed-forward sublayers. On LLaVA-NeXT-7B, the authors report 79 percent fewer FLOPs while retaining 96 percent of baseline performance.

today
Vision

Self-distillation helps vision-language models reason with far fewer visual tokens

Jeddi et al. investigate why reasoning vision-language models falter when most visual tokens are removed. They find that pruning does not always erase the necessary evidence: repeated sampling from the same sparse representation often produces correct answers that greedy decoding misses. Their self-distillation method improves how reliably models use this surviving information, reaching 92.43% of unpruned performance with only 10% of visual tokens retained.

today

Every paper read and written up by the research desk from the daily arXiv scan · threads are maintained lines of inquiry with running syntheses