Encoder-free multimodal models could close the gap with enough compute

Controlled scaling experiments predict that models learning vision directly from pixels can match encoder-based multimodal loss near 10^22 training FLOPs.

Chinese Tech
Lin Chen · Bolin Ni · Qi Yang · Lan Jiang · Kun Ding · Xiaoran Fan · +3 more

CASIA · UCAS · Tencent

Research Digest··3 min read
Chen and colleagues compare multimodal language models that use a pretrained visual encoder with models that instead feed projected image patches directly into the language decoder.

The authors trained two model families using the same sparse decoder ladder, data mixture, optimization setup and visual-token granularity.

Why this paper

From Tencent and 2 others

In one line

Encoder-free multimodal LLMs catch up to encoder-based ones in multimodal loss at roughly 10^22 FLOPs, making them promising at scale.

What we could check

  • ·No code link found
  • ·No weights link found
  • ·No dataset link found
  • ·No compute details found
  • ·No stated limitations found
  • ·No benchmark numbers found

Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.

§
newspaper

Research Digest

Articles published under the Zotpaper byline are synthesized from multiple source publications by our AI editor and reviewed by our editorial process. Each story combines reporting from credible outlets to give readers a balanced, comprehensive view.