How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
- Type
- paper
- Venue
- arXiv
- Year
- 2026
- Source
- arxiv
- Access
- public
- Language
- en
- Added
- 2026-09-29
- Verified
- 2026-09-29
Summary
Studies scaling laws for encoder-free multimodal pretraining (raw visual input into the language model, no separate vision encoder) and asks at what compute budget it catches up with encoder-based models.
Keywords
multimodal · scaling laws · vision encoders · pretraining
Topics
multimodal, scaling laws, vision encoders, pretraining
Research notes
- Discovery: Shared in #random-papers as an arXiv link.
- Method: Fits scaling laws across compute budgets comparing encoder-free vs encoder-based multimodal pretraining.
- Key findings: Encoder-free models underperform at small scale but are projected to catch encoder-based models around 1e22 FLOPs; vision-specific adaptation emerges inside the language model as compute increases.