The authors trained two model families using the same sparse decoder ladder, data mixture, optimization setup and visual-token granularity.
Encoder-free multimodal models could close the gap with enough compute
Controlled scaling experiments predict that models learning vision directly from pixels can match encoder-based multimodal loss near 10^22 training FLOPs.
Chinese Tech
Lin Chen · Bolin Ni · Qi Yang · Lan Jiang · Kun Ding · Xiaoran Fan · +3 more
CASIA · UCAS · Tencent
Research Digest··3 min read
Chen and colleagues compare multimodal language models that use a pretrained visual encoder with models that instead feed projected image patches directly into the language decoder.
Why this paper
From Tencent and 2 others
In one line
Encoder-free multimodal LLMs catch up to encoder-based ones in multimodal loss at roughly 10^22 FLOPs, making them promising at scale.
What we could check
- ·No code link found
- ·No weights link found
- ·No dataset link found
- ·No compute details found
- ·No stated limitations found
- ·No benchmark numbers found
Observed from the paper text and links we have. Absence here means we did not find it, not that it does not exist.
§