Raw video alone, with no captions, can mid-train language models for better vision.

Hwang et al. mid-train Qwen3-1.7B on raw YT-Temporal-1B clips using next-visual-token prediction, gaining 2.9 points on video and 5.1 on image benchmarks while preserving text scores.

Big Tech
Jaedong Hwang · Xiaoqian Shen · Ernie Chang · Changsheng Zhao · Chong Zhou · Saksham Suri · +6 more

Meta AI · Massachusetts Institute of Technology · KAUST · Duke University

Research Digest··2 min read
The authors test whether a pretrained language model can improve its visual understanding by mid-training on raw video frames with no captions or text loss.

Hwang et al.

Why this paper

From Meta AI and 3 others

In one line

Mid-training a pretrained language model on raw video through next-visual-token prediction improves image and video understanding while preserving text performance, without captions.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ✓Limitations stated by the authors (2 noted)
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§

Research Digest

Written by software from the reporting listed above, scored by an automated standards desk, and published without a person reading it first. If something here is wrong, tell the editor and it will be put right.

How we workSubscribe