Vision

Image and video understanding, multimodal perception, detection, segmentation, 3D understanding

30 articles

Vision

Decoupled temporal axis yields motion-prioritized video representations efficiently

The authors conduct a controlled architecture-objective study and propose TT-VidT, which uses a DINOv3-initialized ViT-B/16 per-frame spatial path and a compact Temporal Transfer Layer, trained with Diff Compression to reconstruct frames from a first-frame anchor and motion tokens. They show TT-VidT leads fine-tuning on Jester, Something-Something V2, ARID, and Diving48, improving 54–121% over non-TT methods, with 48–55% fewer encoder FLOPs than DisMo and VideoMAE/V-JEPA2.

3 Oct 2026
Vision

Pruning visual tokens and decoder sublayers cuts multimodal model computation

Chen and colleagues target two sources of multimodal model overhead: redundant visual tokens and unnecessary computation inside the language-model decoder. Their training-free SPIDER framework retains complementary visual information from middle and deep vision layers, then selectively skips attention or feed-forward sublayers. On LLaVA-NeXT-7B, the authors report 79 percent fewer FLOPs while retaining 96 percent of baseline performance.

3 Oct 2026
Vision

Self-distillation helps vision-language models reason with far fewer visual tokens

Jeddi et al. investigate why reasoning vision-language models falter when most visual tokens are removed. They find that pruning does not always erase the necessary evidence: repeated sampling from the same sparse representation often produces correct answers that greedy decoding misses. Their self-distillation method improves how reliably models use this surviving information, reaching 92.43% of unpruned performance with only 10% of visual tokens retained.

3 Oct 2026
Vision

Fine-Grained Rewards Reduce Vision-Language Hallucinations Without Silencing Valid Detail

Long and colleagues address a failure mode of on-policy training for vision-language models: a model can reduce object hallucinations simply by making fewer claims, including fewer correct ones. Their framework instead supplies dense evidence about which objects are present or absent and assigns rewards separately to parts of a response, producing descriptions that the authors report are both more informative and more faithful.

3 Oct 2026
Vision

Steering hidden states stabilizes vision-language reasoning against subtle visual changes

The authors identify answer flips—cases where nearly identical images cause vision-language models to produce different answers—and propose FlipDir, a training-free inference-time method that estimates a low-rank subspace of flip-inducing activations and selectively steers hidden states during decoding. A margin-based gate restricts steering to uncertain steps, recovering original predictions while leaving stable ones unchanged. Across 18 settings spanning scientific reasoning, robot-scene understanding, and medical VQA, FlipDir consistently outperforms existing methods on a combined metric of recovery and preservation.

3 Oct 2026
Vision

Asymmetric certainty gains from optimization hinder multimodal classification

The paper identifies a flaw in multimodal learning: optimization yields asymmetric gains in predictive certainty between strong and weak modalities, driving imbalanced contributions. The authors propose Max Confidence Regularization (MaxCR) to dynamically intervene in modality confidence using a nonlinear sparsity measure, penalizing overconfident strong modalities and encouraging underconfident weak ones. Experiments on several datasets show MaxCR improves overall performance over state-of-the-art baselines.

3 Oct 2026
Vision

Parallel tile inspection helps models find details in large images

Tao et al. address a recurring failure in high-resolution visual question answering: crucial evidence can disappear when an image is resized, yet a model may not know where to zoom from an initial overview. Their Visual Parallel Search framework first examines a grid of tiles concurrently, then lets a main agent combine the tile reports, zoom into selected or merged regions, and answer the question. Across five benchmark splits and three model sizes, this scaffold generally outperformed a sequential zoom-only approach.

1 Oct 2026
Vision

Forensic tool use improves explainable zero-shot image forgery detection

Tan et al. present ATAR, an agentic system that investigates images through repeated visual inspection and calls to specialized forensic tools, rather than explaining a classifier’s predetermined output. Across six zero-shot image manipulation detection and localization benchmarks, ATAR achieved an average image-level F1 score of 78.5%, 11.8 percentage points above the strongest multimodal language model baseline evaluated.

1 Oct 2026
Vision

Visual image rearrangement helps models when exact evidence matters most

Zhang et al. treat multi-image reasoning as a problem of re-representation: reorganising visual evidence so a model can inspect relationships that are difficult to infer from the original images alone. They compare textual reasoning with visual transformations, introduce the grounding-focused MosaicBench, and train MosaicAgent-8B to operate on images. Their central finding is that visual re-representation helps most when answers depend on precise alignment, orientation, or pixel-level comparison, rather than broad semantic interpretation.

1 Oct 2026
Vision

Shared scene graphs connect interaction understanding with grounded robot planning

Zhang and colleagues introduce ChronoGraph, a representation that links actions on functional object parts, such as handles, to subsequent semantic and geometric changes in a scene. They use automatically annotated human videos and simulated robot trajectories to train vision-language models to reconstruct past interactions and predict future states, then demonstrate that the resulting plans can guide a mobile manipulator through existing robot skills.

1 Oct 2026
Vision

Parallel tile inspection and adaptive zoom improve high-resolution visual search

Tao et al. introduce VPS, a visual parallel-search framework in which a main agent delegates tile inspection to parallel question-conditioned sub-agents and then adaptively zooms into relevant regions. Across five benchmark splits and three model sizes, VPS outperforms dedicated zoom-only search in 14 of 15 same-model comparisons, with gains up to 8.0 points. The authors also contribute a supervision pipeline with hint-free verification and a paired role-specific GRPO surrogate for post-training.

30 Sept 2026
Vision

Grounding long-video memories in visual identity lets QA systems track objects across days

The authors introduce Grounded Entity Biographies (GEB), a long-video memory framework that groups visually grounded observations of the same physical instance across clips into retrievable biographies. On four benchmarks, including day-long and week-long recordings, GEB outperforms prior memory frameworks in both multiple-choice and open-ended question answering. On EgoLifeQA it reaches 72.0%, 4.4 percentage points above the best published result.

30 Sept 2026
Vision

Multi-agent framework generates spatial QA data via code execution

Ying et al. introduce Exemplar2VQA, a multi-agent framework that synthesizes large-scale 3D spatial question-answer pairs in simulated environments. By converting user-defined exemplar queries into deterministic code via specialized agents, the system avoids hallucinations common to LLM-based generation. Fine-tuning Qwen2.5-VL models on this synthetic indoor data improves performance on multiple spatial benchmarks, including outdoor and mixed-scene settings.

30 Sept 2026
Vision

End-to-end model writes 3D scene graphs directly from point clouds or Gaussian splats

Milivojevic et al. present GraphWrit3R, a method that takes a 3D point cloud, Gaussian Splats, or a combination as input and directly outputs a complete scene graph as a structured JSON script. On the 3DSSG benchmark, the method achieves state-of-the-art performance on object class, predicate, and triplet recall, outperforming methods that rely on ground-truth object annotations during inference.

29 Sept 2026
Vision

Branched YOLOv2 with geometric features improves traffic sign recognition speed and accuracy.

The authors extended YOLOv2 with intermediate prediction layers (branched architecture) to allow early exit for easy cases, and introduced geometric features from Bayesian segmentation to distinguish visually similar signs. On a combined GTSDB/GTSRB dataset of ten classes, branched YOLOv2 achieved similar accuracy but slightly faster runtime (0.647 s vs 0.6607 s), and adding geometric verification during inference improved mAP from 0.680 to 0.713.

22 Sept 2026
Vision

Streaming video memory works better when internalized as evolving latent tokens

The authors introduce LatentStream, a progressive latent working-memory framework for multimodal LLMs processing streaming video. Instead of storing historical clips in an external memory bank and retrieving them as extra visual context, LatentStream converts retrieved evidence into compact, fixed-length latent memory tokens that evolve over time. Experiments show the method outperforms prior store-and-retrieve approaches on both online and offline video understanding benchmarks.

7 Sept 2026