Skip to content
AI.info

Research

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining

Scaling Laws for Encoder-Free Multimodal Pretraining Overview Research area: Computer vision and multimodal large language models (MLLMs) — specifically the comparison between architectures that use a

How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
arXiv
2609.35457
Published
2026-09-28
Authors
Lin Chen, Bolin Ni, Qi Yang, Lan Jiang, Kun Ding, Xiaoran Fan, Hower Yang, Ying Wang, Shiming Xiang

AI summary

Scaling Laws for Encoder-Free Multimodal Pretraining

Overview

Research area: Computer vision and multimodal large language models (MLLMs) — specifically the comparison between architectures that use a pretrained visual encoder and "encoder-free" architectures that learn visual representations directly from raw pixels, characterized through scaling laws.

Technical level: Intermediate. The paper assumes familiarity with transformer decoders, Mixture-of-Experts (MoE) layers, attention masking, and the Chinchilla-style IsoFLOP scaling-law methodology.

Scope: A controlled scaling study on a matched ladder of 11 sparse MoE decoders (1.1B–44B total parameters) that fits separate scaling laws for text and multimodal objectives in order to quantify how the efficiency gap between encoder-free and encoder-based MLLMs evolves with compute.

What This Paper Is About

Most multimodal LLMs inherit a strong visual prior from a pretrained visual encoder such as a SigLIP 2 ViT. Encoder-free models remove that component and feed projected image patches straight into the language decoder, which must then learn visual representations from raw pixels. The paper asks how far this design is from parity: it systematically measures whether encoder-free models can catch up with encoder-based models as compute grows, and if so, at what training budget — and it examines what the decoder does internally to compensate for the missing encoder.

Key Contributions

  1. A controlled scaling-law comparison of encoder-free and encoder-based MLLMs. The two families share the same sparse decoder ladder of 11 MoE models (1.1B–44B total, 71M–2.4B active non-embedding parameters), the same data mixture, optimization setup, and visual-token granularity. Separate laws are fitted for $\mathcal{L}{\mathrm{text}}$ and $\mathcal{L}{\mathrm{mm}}$.

  2. A prediction of when encoder-free models catch up. Using the fitted laws, the authors extrapolate the compute budget at which the multimodal loss frontier of encoder-free models crosses that of encoder-based models, both under compute-optimal allocation and under overtraining.

  3. A topic-level breakdown of the crossover. The aggregate multimodal loss is decomposed into topics (STEM, Charts, GUI, OCR, Caption, and others), and the crossover point is estimated separately for each.

  4. A mechanistic probe of how the decoder absorbs visual encoding. Three internal probes — attention over visual tokens, layerwise representation drift, and MoE expert routing imbalance (MaxVio) — characterize the "vision-specific adaptation" of the encoder-free decoder.

Main Findings

  • Compute-optimal allocation shifts toward larger models for vision. On the text objective, the model allocation exponent is nearly identical between the two architectures ($a = 0.427$ for encoder-free versus $a = 0.422$ for encoder-based). On the multimodal objective, removing the encoder raises $a$ from $0.464$ to $0.570$, with bootstrap 80% intervals of $[0.546, 0.595]$ and $[0.458, 0.472]$ respectively. The shift persists under causal attention over visual tokens ($a = 0.557$).

  • Text frontiers nearly overlap. The fitted text loss–compute exponents are $0.0973$ (encoder-free) and $0.0979$ (encoder-based). Measured and fitted $\mathrm{EG}^{C}$ stay around $0.98$ and $\mathrm{EG}^{M}$ around $0.99$. The authors describe text as a "nearly matched control."

  • Multimodal frontier diverges but converges with scale. Encoder-free models require more compute for equal multimodal loss throughout the measured range, but their loss falls faster with compute. The fitted loss–compute exponents are $0.3778$ versus $0.2998$ (envelope estimator: $0.3668$ versus $0.3050$). At equal loss, encoder-free models favor a larger decoder ($\mathrm{EG}^{M} \approx 0.80$, about $1.25\times$ the FLOPs per token).

  • Predicted crossover on the order of $10^{22}$ FLOPs. The point estimate is $6.1 \times 10^{21}$ FLOPs under compute-optimal allocation, with a conditional bootstrap 80% interval of $[4.2 \times 10^{21}, 1.0 \times 10^{22}]$. For reference, the paper estimates the pretraining compute of Kimi K2.5 at approximately $10^{25}$ FLOPs, using $C \approx 6N_{\mathrm{active}}D$ — a gap of roughly three orders of magnitude.

  • Overtraining delays the crossover. Under overtraining at $k = 5$, the point estimate is $1.2 \times 10^{22}$ FLOPs with an 80% interval of $[8.4 \times 10^{21}, 2.0 \times 10^{22}]$. At the largest fitted budget, overtraining lowers $\mathrm{EG}^{C}$ from $0.62$ to $0.52$ and $\mathrm{EG}^{M}$ from $0.80$ to $0.74$. The authors connect this to the allocation shift: since encoder-free models favor larger decoders on multimodal data, spending extra compute on tokens at fixed model scale benefits them less. On text, both ratios change by less than 1% from their $k = 1$ values.

  • Crossover ordering follows reliance on visual priors. Encoder-free models catch up first on STEM (already close to parity at the largest fitted budget), then Charts, and considerably later on GUI, OCR, and Caption. STEM is described as consisting mainly of text, symbols, and simple diagrams; captioning and GUI/OCR require rich natural-image and spatial perception.

  • Bidirectional attention among visual tokens grows more valuable with scale. Comparing against a fully causal variant: causal attention is slightly better on text at measured budgets ($\mathrm{EG}^{C} = 1.011$, reaching equal loss for about 1% less compute) but the gain shrinks with compute. On multimodal, causal is mildly worse on average ($\mathrm{EG}^{C} = 0.990$), and the multimodal benefit of bidirectional attention increases with compute.

  • Shallow decoder layers act as an implicit visual encoding stage. Cosine similarity to layer-0 inputs shows visual tokens in encoder-free models diverge from their inputs much earlier than in encoder-based models, while text token trajectories remain close across both systems. The pattern holds across model scales.

  • Expert routing concentrates for visual tokens. Using MaxVio (relative excess of the most-loaded expert over perfectly balanced load), the two architectures are similarly low when visual and text tokens are aggregated. Separated by modality, both show greater imbalance; for text tokens the architectures remain closely matched, while across the four largest model sizes encoder-free models show consistently higher average MaxVio and a wider band for visual tokens.

  • A sharp loss drop marks the decoder bootstrapping visual representations. Encoder-based models decrease smoothly; encoder-free models decrease slowly and then drop sharply. In the 8B models at layer 12, attention to visual tokens rises from $0.217$ to $0.645$ across this drop, approaching the encoder-based level. At every multimodal IsoFLOP budget the compute-optimal models have already passed this drop.

  • Extrapolation checks support the fits. Held-out forecasting errors are $-1.03%$ and $-0.87%$ on text (a $5\times$ extrapolation to $4 \times 10^{20}$ FLOPs) and $-0.01%$ and $+2.19%$ on multimodal (a $2.5\times$ extrapolation to $1 \times 10^{21}$ FLOPs) — within about 2% overall.

Methodology in Plain English

The authors build two model families that are as comparable as possible, varying only in how images enter the decoder.

  • Encoder-based arm: images go through a pretrained SigLIP 2 ViT (27 layers, width 1152, patch size 16, AnyRes), then a $2\times2$ ConvPool adapter and a projector. The ViT is about 400M parameters and is trained jointly with the decoder, with its size held fixed across the ladder.
  • Encoder-free arm: raw image patches go straight into the decoder through a patch projection, using the same $2\times2$ merging so both arms produce one visual token per $32\times32$ pixel region and the same token count. Visual tokens attend bidirectionally within each image; all other attention stays causal. A fully causal variant is also studied.

Both arms share a ladder of 11 sparse MoE decoders. Every rung uses 256 routed experts, activates the top 8 per token plus one shared expert, so the activation ratio is fixed at $8/256 = 1/32$. Training uses sequence length 4,096, the Muon optimizer, 2,000 warmup steps, and a 1:1 text-to-multimodal data mixture. Batch size and learning rate follow an internal scaling law; identical hyperparameters are used for both arms at each rung.

To estimate scaling behavior, the researchers use the Chinchilla-style IsoFLOP procedure: at each compute budget $C$, they vary model size $M$, set $D = C/M$, fit validation loss as a quadratic in $\log M$, and take the vertex as $M_{\mathrm{opt}}(C)$. They use six logarithmically spaced budgets per objective — $2 \times 10^{19}$ to $2 \times 10^{20}$ FLOPs for text and $1 \times 10^{20}$ to $1 \times 10^{21}$ FLOPs for multimodal. Fitting $M_{\mathrm{opt}} \propto C^{a}$ and $D_{\mathrm{opt}} \propto C^{b}$ (with $a+b=1$) gives the allocation law; fitting $\mathcal{L}^{*}(C) = E + KC^{-\gamma}$ gives the loss frontier.

Efficiency is reported as two ratios following MAI-Thinking-1: $\mathrm{EG}^{C}$ (reference compute over target compute at equal loss) and $\mathrm{EG}^{M}$ (reference FLOPs per token over target FLOPs per token at equal loss), with encoder-free as target and encoder-based as reference. Values below 1.0 mean encoder-free needs more compute or a larger decoder.

Overtraining is modeled as a prefactor shift in the loss law, $\mathcal{L}(C_{\mathrm{base}}, k) = E + g(k)KC_{\mathrm{base}}^{-\gamma}$, where $E$ and $\gamma$ are reused from the compute-optimal fit and only $g(k)$ is estimated empirically at $k \in {2,3,4,5}$. Robustness checks include a validation-curve envelope estimator, held-out extrapolation tests, and a residual bootstrap for uncertainty.

Why This Matters

Impact on research. The paper reframes the encoder-free question from "does it work?" to "at what compute does it become competitive?" By fitting separate laws for text and multimodal objectives on a shared ladder, it isolates the visual prior as the variable of interest rather than confounding it with data mixture or optimization differences — something earlier encoder-free studies (Fuyu, EVE, SOLO, SAIL) did not do. It also gives a mechanistic account of where the missing encoder's function goes, which is directly actionable for architecture design rather than being a purely empirical feasibility claim.

Real-world applications (the paper reports topic-level results for these task families):

  • Document and chart understanding — Charts narrows quickly and is the second topic to reach crossover, relevant to analytics and reporting tools.
  • STEM and technical content — nearly at parity at the largest fitted budget, relevant to scientific and educational assistants.
  • GUI agents and OCR — these cross over much later in the fitted extrapolation, meaning encoder-free designs are currently a worse fit for screen-reading and interface automation at these budgets.
  • Image captioning — captioning is the latest to converge, indicating encoder-free models are least competitive where rich natural-image description is required.

Industry relevance. The Kimi K2.5 comparison ($\approx 10^{25}$ FLOPs) matters because it places the predicted $10^{22}$-FLOP crossover roughly three orders of magnitude below the compute used by recent flagship models. If the fitted laws hold, an encoder-free architecture could be a viable choice inside practical pretraining budgets — and it removes an entire pretrained component from the stack, simplifying the pipeline and unifying the architecture around a single transformer.

Future Directions

  1. Designing decoders for native visual representation learning. The authors argue the vision-specific adaptations observed — bidirectional visual attention, early-layer visual processing, concentrated expert routing — suggest encoder-free models may need decoder architectures designed explicitly for native vision rather than inherited language designs.

  2. Jointly scaling the visual encoder and the decoder. The current study holds the ViT at a fixed ~400M scale, so its reported trends are conditional on that regime. The paper explicitly notes that jointly scaling the encoder would define a different allocation problem requiring a separate sweep balancing representation gains against front-end compute — and cites prior work showing coupled scaling between visual encoder and language model under data constraints.

  3. Accounting for visual encoder FLOPs. The main analysis uses decoder FLOPs as the primary compute measure. Appendix C reports that including visual encoder FLOPs leaves the main conclusions unchanged and further strengthens the relative efficiency of encoder-free models, but a fuller accounting could be pursued.

  4. Improving compute efficiency of native visual input. The conclusion explicitly points toward training strategies that speed up the early "bootstrapping" phase, during which visual tokens are uninformative and largely ignored before the sharp loss drop.

Target Audience

Researchers and engineers working on multimodal large language model architectures and pretraining, particularly those evaluating whether to keep or remove a pretrained vision encoder. It is also relevant to practitioners of scaling-law methodology who want an example of applying IsoFLOP analysis to a multimodal objective, and to teams making compute-allocation decisions about model size versus training tokens for vision-language models. Readers should be comfortable interpreting allocation exponents, loss–compute exponents, and MoE routing statistics.

Authors’ abstract

Most modern multimodal large language models (MLLMs) build on a pretrained visual encoder that provides a strong visual prior. Encoder-free MLLMs instead learn visual representations directly from raw pixels, offering a simple and unified architecture, but their scaling behavior has not been systematically characterized. To fill this gap, we compare scaling laws for encoder-free and encoder-based MLLMs and report three main findings: (1) Removing the visual encoder shifts the compute-optimal allocation for the multimodal objective toward larger models, while leaving that for text nearly unchanged. (2) The two architectures exhibit nearly overlapping loss--compute frontiers on the text objective, but diverge on the multimodal objective: encoder-free models underperform at small scales yet are predicted to catch up at around $10^{22}$ FLOPs, well within practical pretraining budgets. (3) Without a visual encoder, the language model learns to take over its role via vision-specific adaptation: bidirectional interactions among visual tokens become increasingly beneficial as training compute grows, visual processing shifts toward earlier layers, and expert routing for visual tokens becomes more concentrated. Overall, our results indicate that the advantage of the visual prior provided by a pretrained encoder diminishes with scale, positioning encoder-free architectures as a promising direction for multimodal pretraining.

Read the original paper