← Back to explorer

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Type
paper
Venue
arXiv
Year
2026
Source
arxiv
Access
public
Language
en
Added
2026-09-29
Verified
2026-09-29

Summary

Studies scaling laws for encoder-free multimodal pretraining (raw visual input into the language model, no separate vision encoder) and asks at what compute budget it catches up with encoder-based models.

Keywords

multimodal · scaling laws · vision encoders · pretraining

Topics

multimodal, scaling laws, vision encoders, pretraining

Research notes

  • Discovery: Shared in #random-papers as an arXiv link.
  • Method: Fits scaling laws across compute budgets comparing encoder-free vs encoder-based multimodal pretraining.
  • Key findings: Encoder-free models underperform at small scale but are projected to catch encoder-based models around 1e22 FLOPs; vision-specific adaptation emerges inside the language model as compute increases.