Research
Mixture of Layers: Dynamic Layer Routing for Visual Reasoning
Overview Research area: Multimodal large language models (MLLMs) — specifically, how visual features are selected from intermediate layers of pre-trained vision encoders for fine-grained visual reason

- arXiv
- 2610.09440
- Published
- 2026-10-07
- Authors
- Jeonghwan Kim, Sofia Stoica, Jiwan Chung, Ansel Blume, Hyeonjeong Ha, Zhenhailong Wang, Xin Luna Dong, Heng Ji
AI summary
Overview
Research area: Multimodal large language models (MLLMs) — specifically, how visual features are selected from intermediate layers of pre-trained vision encoders for fine-grained visual reasoning (arXiv:2610.09440v1 [cs.CV], 07 Oct 2026, CC BY-SA 4.0).
Technical level: Advanced. The paper assumes familiarity with vision-language architectures, mixture-of-experts routing, sparse top-k selection, softmax/KL-divergence math, and MLLM training pipelines.
Scope: The paper proposes Mixture of Layers (MoL), an instruction-conditioned routing mechanism that dynamically selects which intermediate vision-encoder layers feed the language model, and evaluates it across 7 fine-grained visual reasoning benchmarks on multiple backbones and vision encoders.
What This Paper Is About
Most MLLMs feed the language model only the final or penultimate vision-encoder representations (or apply fixed aggregation rules), which makes the visual abstraction the same regardless of what the user asks. That is a problem for fine-grained tasks where the answer depends on small objects, localized text, counting, subtle attributes, or spatial relations that a single late-layer representation may not preserve. The goal of this paper is to let the model choose, per instruction — and optionally per image patch — which vision encoder layers to read from, without adding more visual tokens or extra vision encoders.
Key Contributions
- Instruction-conditioned layer routing formulation. MoL adaptively selects intermediate vision-encoder layer representations conditioned on the input instruction, supporting both image-level routing (MoL_layer) and per-patch routing (MoL_patch) as sparse top-k aggregation over selected hidden states.
- Hybrid routing with a disagreement-gated reserve layer. MoL_hybrid combines layer-level and patch-level distributions via a weighted product-of-experts and uses their disagreement (KL divergence) to gate a reserve layer, balancing local fine-grained evidence with global semantic context.
- Empirical gains without extra tokens or extra encoders. Across seven fine-grained visual reasoning benchmarks, MoL improves fine-grained visual reasoning without increasing the number of visual tokens passed into the LLM and without relying on multiple vision encoders, making it complementary to existing MLLM scaling approaches.
- Analysis of why layer-wise sampling helps. The paper studies vision encoders' receptive field scales across layers and their sampling behaviors, arguing that conditional visual representations are a key step toward better visual perception and reasoning in MLLMs.
Main Findings
- Large gains on fine-grained benchmarks. The abstract reports +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to baseline MLLMs, without multi-resolution inputs, interleaving of multiple vision encoders, or increasing patch-token counts.
- Gains hold across backbones and encoders. Table 1 shows best absolute improvements over each backbone baseline: LLaVA-v1.5-13B (Vicuna-13B + CLIP) +13.44 on V* all; Vicuna-13B + DINOv2 w/ Txt +18.90 on V* all; Llama-3-8B-V +8.42 on V* all; Phi-1.5-1.3B + SigLIP +8.83 on V* all. On Phi-1.5-1.3B, MoL improves OCR by +16.67 and CharXiv by +16.39.
- Different variants excel on different task types. MoL_patch is often strongest on localized perceptual categories (e.g., V* all 84.87 on LLaVA-v1.5-13B; OCR 60.00; direct attribute 86.09), while MoL_hybrid tends to do better on benchmarks requiring both local grounding and broader reasoning (e.g., GPT4V-hard 64.71, HRBench8K 38.62).
- Beats prior multi-layer feature aggregation. In Table 2, MoL_patch reaches the best overall V* score (84.87), outperforming Dense Connector (81.35) by 3.52 points, with strong localized-category gains. MoL_hybrid improves on harder benchmarks such as GPT4V-hard, MMStar, and HRBench.
- Single-encoder MoL competes with multi-encoder interleaving. Table 3 shows a dual-encoder MoL_hybrid configuration (CLIP + DINOv2 w/ Txt, Vicuna-13B) reaching 90.34 on V* versus 82.35 for Interleaved-MoF, with the largest gap on GPT4V-hard (82.35 versus 58.82), then OCR (66.67 versus 53.33) and direct attributes (92.17 versus 82.61); relative-position accuracy is tied at 98.68. The paper states this corresponds to 19 additional correct answers out of 238 (four on GPT4V-hard, four on OCR, eleven on direct attributes), and that the remaining comparisons are mixed, including lower HRBench4K and RealWorldQA scores.
- The authors explicitly decline to claim a uniform multi-encoder win. They state that the comparison is between complete systems, not a controlled routing-only ablation, and that a matched no-router model with the same fusion, training schedule, and checkpoint-selection rule is needed to isolate routing from fusion and encoder choice.
- Coarse-grained performance is largely preserved. Table 4 shows small changes on standard VQA: on LLaVA-v1.5-13B, best deltas are POPE +0.87, GQA +0.49, TextVQA -0.14, MMMU +1.33; on Vicuna-13B + DINOv2 w/ Txt they are POPE +0.07, GQA -0.89, TextVQA +0.25, MMMU +0.00.
- Routing is cheap. The router adds negligible parameter overhead, stated as <0.02% over the 13B base model, and token-by-token language decoding is left unchanged.
- Router behavior differs systematically by variant. The layer-level router (MoL_layer) consistently prioritizes early layers, especially for localized categories such as OCR, direct attributes, and relative position; the patch-level router (MoL_patch) places more mass on later layers, and MoL_hybrid yields more consistent mid-to-late layer usage. Receptive field increases with depth — early layers capture local interactions while later layers aggregate broader spatial context.
- Note on reported numbers. The CharXiv scores listed for MoL_layer, MoL_patch, and MoL_hybrid on LLaVA-v1.5-13B differ between Table 1 (26.25, 29.17, 25.56) and Table 2 (27.85, 28.65, 26.97) in the supplied content.
Methodology in Plain English
The setup is a standard MLLM: one or more pre-trained vision encoders (CLIP, DINOv2 w/ Txt, SigLIP), a connector that projects visual features into the language model's space, and an LLM backbone (Vicuna-13B / Llama-3-8B-Instruct / Phi-1.5-1.3B). Instead of taking only the last or penultimate layer, the authors keep all L layer representations available and treat each layer as an "expert," similar to mixture-of-experts.
A text prompt is embedded by the vision tower's aligned text encoder and pooled into a single query vector. That query is compared against each layer's patch features to produce scores, which become routing probabilities. Three variants use those probabilities differently: MoL_layer pools each layer over the whole image and picks one shared set of top-k layers; MoL_patch makes a separate top-k choice for every patch, so different image regions can draw on different abstraction levels; MoL_hybrid mixes a global distribution and a patch-level distribution with a learned, prompt-dependent weight (a weighted product-of-experts), and measures their disagreement to decide how much to fall back on a designated "reserve layer" (a last or penultimate layer). The selected representations are combined, normalized, and passed into the LLM as before — the number of output patch tokens does not change.
Training uses the standard autoregressive language-modeling loss plus an auxiliary load-balancing loss taken from Switch transformers to prevent the router from collapsing onto a few layers, with a coefficient controlling its strength. For multiple encoders, each encoder is routed independently and the routed streams are projected, upsampled to a common token grid (DINOv2 and CLIP produce 256 and 576 patch tokens respectively in their setting), concatenated, and passed through a learned linear projection.
Evaluation uses the exact pre-training and instruction-tuning datasets of each backbone for fair comparison, with fine-grained benchmarks V*, HRBench4K, HRBench8K, NaturalBench, RealWorldQA, CharXiv, and MMStar.
Why This Matters
Impact on research. The paper reframes visual feature selection as an instruction-conditioned routing problem rather than a fixed architectural choice. It provides evidence that intermediate vision layers carry query-relevant information that later layers do not preserve, and it does so without the usual remedies of more tokens, more resolution, or more encoders — which matters because those remedies are expensive and often confound each other. It also leaves a clear methodological marker by stating that its multi-encoder result is not a controlled ablation.
Real-world applications.
- Document and diagram understanding, where answers depend on localized text and layout details (the paper cites document/diagram understanding and CharXiv-style chart reasoning).
- Industrial inspection, where small defects and subtle visual attributes matter (named in the paper as a motivating real-world task).
- OCR-heavy and text-in-image workflows, where MoL shows large gains (e.g., +10.00 and +16.67 OCR improvements in Table 1 settings).
- Counting, spatial-relation, and attribute-recognition queries in assistants, where the paper reports gains such as +20.87 on direct attribute recognition and +28.94 on relative position.
Industry relevance. Because the router adds <0.02% parameter overhead over a 13B base model and leaves language decoding unchanged, MoL is an incremental add-on to existing MLLM pipelines rather than a re-architecture. It can be layered on top of existing high-resolution and multi-patch designs, and it avoids the inference cost of feeding more visual tokens or running multiple vision encoders.
Future Directions
- A properly controlled routing ablation. The authors state that a matched no-router model with the same fusion, training schedule, and checkpoint-selection rule is needed to separate the contribution of routing from fusion and encoder choice.
- Integrating with resolution and token-budget scaling. The paper positions MoL as complementary to existing multi-patch and high-resolution approaches and to multi-encoder stacking, which suggests testing combinations rather than replacements.
- Understanding routing decisions more deeply. The analysis of receptive fields, layer selection patterns, and patch-level sampling behavior is presented as a first step; extending it (including to DINOv2 variants, deferred to the appendix) is an open line of work.
- Broadening backbones, encoders, and layer counts. Results span 1.5B, 8B, and 13B backbones and three vision encoders; the paper does not report whether the recipe transfers further, and the partial setting (k=1 for the multi-encoder configuration, k=4 in the layer-selection figure) leaves room to study routing sparsity.
Target Audience
Researchers and engineers working on multimodal large language models, vision-language connectors, and mixture-of-experts routing, especially those focused on fine-grained visual perception, OCR, document and chart understanding, or high-resolution visual reasoning. It is also useful for practitioners deciding between adding visual tokens, adding vision encoders, or changing how existing encoder features are selected, and for anyone studying how information is distributed across vision-encoder depth.
Authors’ abstract
Pre-trained vision encoders contain layer-wise visual representations that differ in spatial granularity, semantic abstraction, and sensitivity to local details. However, most Multimodal Large Language Models (MLLMs) rely on only the final or penultimate vision encoder representations or fixed aggregation rules, making visual abstraction largely query-agnostic and limiting access to fine-grained cues such as small objects, spatial details, text, and subtle visual attributes. In this work, we propose Mixture of Layers (MoL), an instruction-conditioned layer routing approach at the visual patch level that dynamically aggregates query-relevant latent representations from intermediate vision encoder layers. Given a text query, MoL predicts routing probabilities over vision encoder layers and performs a top-k sparse aggregation over selected hidden states at either the image level, patch level, or through a hybrid routing mechanism. In doing so, MoL enables query-adaptive access to layer-specific visual features for fine-grained visual reasoning. Our experiments across 7 fine-grained visual reasoning tasks demonstrate substantial performance improvements, especially across fine-grained visual grounding and understanding tasks such as +18.9% improvement on V* in overall accuracy, +4.5% on HRBench4K, and +16.3% on CharXiv compared to the baseline MLLMs, without resorting to multi-resolution inputs, simple interleaving of multiple vision encoders, or increasing the number of patch tokens. We study vision encoders' receptive field scales across different layers and their sampling behaviors to provide an in-depth analysis of why layer-wise sampling is helpful, demonstrating that conditional visual representations are a key step towards better visual perception and reasoning in MLLMs. Our project page is available at https://wjdghks950.github.io/mol.github.io/.