Hwang et al.
Raw video alone, with no captions, can mid-train language models for better vision.
Hwang et al. mid-train Qwen3-1.7B on raw YT-Temporal-1B clips using next-visual-token prediction, gaining 2.9 points on video and 5.1 on image benchmarks while preserving text scores.
Big Tech
Jaedong Hwang · Xiaoqian Shen · Ernie Chang · Changsheng Zhao · Chong Zhou · Saksham Suri · +6 more
Meta AI · Massachusetts Institute of Technology · KAUST · Duke University
Research Digest··2 min read
The authors test whether a pretrained language model can improve its visual understanding by mid-training on raw video frames with no captions or text loss.
Why this paper
From Meta AI and 3 others
In one line
Mid-training a pretrained language model on raw video through next-visual-token prediction improves image and video understanding while preserving text performance, without captions.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ✓Limitations stated by the authors (2 noted)
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§