Skip to content
AI.info

Research

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs

Overview Research area: Model merging and Mixture-of-Experts (MoE) large language models; specifically the interpretation and diagnosis of routing changes after merging. Technical level: Advanced. The

Routing Drift Alone Does Not Diagnose Failure in Merged MoE LLMs
arXiv
2609.32821
Published
2026-09-26
Authors
Yuanyi Wang, Yanggan Gu, Su Lu, Guanghao Zhu, Pengkai Wang, Yifan Yang, Congkai Xie, Zhaoyi Yan, Jianmin Wu, Hongxia Yang

AI summary

Overview

Research area: Model merging and Mixture-of-Experts (MoE) large language models; specifically the interpretation and diagnosis of routing changes after merging.

Technical level: Advanced. The paper assumes familiarity with sparse MoE architectures, top-k routing, router logits, model merging (Task Arithmetic, TIES), and intervention-based causal evaluation.

Scope: A diagnostic study across three MoE LLM families testing whether routing drift after merging is evidence of routing failure, plus a case-study repair method called Selective Router Repair (SRR).

What This Paper Is About

When specialized MoE LLMs are merged without joint retraining, tokens often get sent to different experts than they would in the source models. This "routing drift" is commonly read as evidence that the router has broken, motivating recent work on routing realignment. This paper asks whether routing drift actually indicates routing failure — degradation in task-relevant behavior — and what evidence should justify repairing a merged router.

Key Contributions

  1. Routing drift analysis. Controlled counterfactual interventions that cross source and merged router inputs and parameters on identical token sequences across DeepSeekMoE, OLMoE, and Qwen3-MoE under Average and Task Arithmetic merging. Within changed-route events, 77.7–96.9% are attributed to input (representation) shifts, while router-parameter-induced changes account for at most 1.8%.

  2. Task-grounded diagnosis. A formalization of routing failure as intervention-relative recoverable task loss — the expected task-loss difference between a baseline routing policy and a specified alternative policy, with non-routing parameters held fixed — implemented as a paired routing test toolkit.

  3. Repair supervision case study (SRR). Selective Router Repair fits likelihood-weighted, source-derived expert-pair corrections on merged router inputs, updating only selected router rows in the final five sparse layers. It is used to test whether source-likelihood advantages identify beneficial local corrections.

  4. Negative and inconclusive results reported as such. The paper reports that source-route restoration and fitted SRR updates do not establish reliable task benefits on the tested checkpoints, while a deliberately corrupted-router control does show recoverable loss.

Main Findings

  • Route changes are mostly representation-induced, not router-parameter-induced. Across six settings (three architectures × two merging methods), using 256 domain-balanced prompts per setting and all continuation tokens and sparse layers, 26.5–54.3% of token–layer expert sets change. Among changed routes, 77.7–96.9% change under representation-only replacement but not router-only replacement; router-parameter-induced changes account for at most 1.8%.

  • Structural routing differences poorly predict intervention gains. Positive and negative source-route gains occur at overlapping full-distribution Jensen–Shannon divergences across all three architectures, and JS-based prediction on the final five sparse layers remains near chance (AUROC 0.47–0.52; chance 0.5).

  • Source-route restoration shows no reliable task improvement. Replacing routes throughout question–choice sequences and scoring answer tokens gives correct-choice-margin changes whose 95% confidence intervals span zero. Local tests on DeepSeekMoE and OLMoE show the merged route is rarely best within fixed candidate sets, yet the source route wins only roughly half of the comparisons.

  • Different routes can produce similar mixture outputs. At a fixed merged hidden state with expert parameters fixed, the mean maximum output cosine between entering and leaving experts is 0.041–0.096, while source-route and native merged-route mixtures have mean cosine 0.896–0.976. Mixture cosine exceeds the expert-pair maximum in all 727 sampled events. Observed routes exceed 32 matched random-route controls (matched on expert count, overlap with the merged route, and routing-weight values) in mean cosine by 0.112–0.147, with positive paired 95% prompt-bootstrap intervals in every setting.

  • A positive control confirms the diagnostic can detect recoverable loss. Permuting OLMoE router logits in nested sets of k ∈ {1, 4, 16} layers (two fixed randomizations per Average and Task Arithmetic parent, 256 fixed ARC items per parent) yields 12.50–14.84 percentage points of recovery from clean-route replay at k = 4 and 26.56–32.81 pp at k = 16. All eight comparisons at k = 4, 16 pass Holm correction across the twelve tests; none at k = 1 does. Replay accuracy ranges from 50.39% to 51.56% around the 50.78% clean reference.

  • Natural routing alternatives do not establish recovery. For source-route replay and a frozen LC-calibrated route schedule, neither accuracy nor correct-choice margin has a strictly positive 95% interval in any displayed setting. LC replay changes correctness on just one item per parent: −0.39 pp for Average and +0.39 pp for TA.

  • Source-informed directions agree in sign but do not deliver utility. Of 1024 diagnostic events (256 prompts each for OLMoE and Qwen3-MoE under Average and TA candidates), 471 have positive source advantage. Source and fitted directions agree in sign on 83.4% of supported events (95% CI [80.0, 86.6]%), yet neither pooled utility nor opposite-direction contrast establishes an improvement. On 440 matched pairs, the native contrast is −0.341 (95% CI [−3.323, 2.631]) and the opposite contrast is −0.997 (95% CI [−5.653, 3.868]) in 10⁻³ nats/token.

  • Fitted corrections do not execute as predicted at later layers. On 128 prompts per setting across eight OLMoE/Qwen3-MoE Parent/SRR pairs (5,120 events total), the input-response RMS at the final updated layer is approximately 0.38–0.86 times the direct-correction RMS.

  • Benchmark effects are small and mixed. Using MMLU, HellaSwag, ARC-Challenge, ARC-Easy, PIQA, WinoGrande, BoolQ, and GSM8K with five-shot lm-evaluation-harness and a VLLM backend, over 35,326 shared item identities per pair evaluated five times, five of the eight OLMoE/Qwen3-MoE settings show positive average-score changes, ranging overall from −0.056 to +0.139 pp. Against HARC, each method has the higher point estimate in four cases. On OLMoE–WUDI, SRR (−0.021 pp) exceeds HARC (−0.075 pp) yet both fall below Parent. On DeepSeekMoE, HARC ranges from −0.133 pp under TIES to +0.236 pp under Average merging.

Methodology in Plain English

The authors first build a toolkit that lets them swap components of a routing decision independently. For a source model and a merged model, they form four combinations: source router with source hidden state, merged router with source hidden state, source router with merged hidden state, and merged router with merged hidden state. By conditioning on tokens where the source and merged routes disagree, they can ask whether the disagreement came from the input representation, the router parameters, or both.

Second, they check whether those disagreements predict anything useful. They measure routing differences with metrics such as Jensen–Shannon divergence and top-k set distance, then see whether those numbers predict the next-token likelihood gain from restoring the source route. They also restore source routes on full question–choice sequences and compare correct-choice margins.

Third, they inspect what happens inside a MoE layer when routes differ. Holding the merged hidden state and expert parameters fixed, they compare the mixture output produced by the source route against the native merged route, and against random routes matched on expert count, route overlap, and weight values.

Fourth, they define routing failure operationally. Failure means task loss that can be recovered by switching from one routing policy to a specified alternative, with all non-routing parameters frozen. They validate this with a positive control that deliberately corrupts router logits and then replays clean routes, and apply it to natural alternatives such as source-route replay and a frozen calibrated schedule.

Finally, they construct SRR as a case study. SRR computes a weighted source-minus-parent preference profile over experts, screens for expert pairs with sufficient activity (≥ 0.002) and sign consistency (≥ 0.75), then fits a ridge-style regression for each pair on merged router inputs to predict clipped logit residuals. Selected router rows are updated in opposite directions by a scaled half-step. Hyperparameters are τ = 0.5, γ = 10⁻³, η = 0.0625, solved with preconditioned conjugate gradients. The authors then test the local direction of these updates, whether the updates execute as predicted during joint inference, and whether they change benchmark scores relative to the unrepaired parent.

Why This Matters

Impact on research. The paper separates a candidate construction procedure from demonstrated recovery. It argues that source-model agreement is a repair objective, not a label of failure, and that post-merge routing papers should justify repairs with task-level intervention effects rather than routing mismatch metrics. It also supplies a positive control showing that the proposed criterion can detect recoverable loss when loss is deliberately imposed.

Real-world applications.

  • Model-merging pipelines: teams that combine domain-specialized MoE checkpoints can avoid spending effort "fixing" routing that is not causing measurable harm.
  • Deployment and evaluation gating: the intervention-relative recoverable loss criterion can serve as a release check for whether a merged checkpoint needs router repair at all.
  • Compression and expert pruning: the finding that different expert selections can yield highly similar mixture outputs (mean cosine 0.896–0.976) is relevant to selecting which experts to keep.
  • Multi-model serving infrastructure: a single merged MoE is cheaper to serve than separate specialists; knowing when merging preserves useful behavior informs whether a merged deployment is justified.

Industry relevance. Merging avoids joint retraining and ensemble inference, both of which are costly at LLM scale. The paper reports experiments on DeepSeekMoE-16B-A3B, OLMoE-7B-A1B, and Qwen3-30B-A3B across four merging methods (Average, Task Arithmetic, TIES, WUDI-Merge), on 8 A800 GPUs, which is a realistic production-adjacent scale.

Future Directions

  • Other routing interventions. The authors state that their result does not exclude task loss recoverable under routing policies other than the ones they tested. Causal, calibrated, or learned router swaps remain open.
  • Attributing degradation to router parameters. The paper notes that recovery through routing does not by itself identify router-parameter changes as the original cause of degradation, leaving the causal chain from merge to harm unresolved.
  • Improving repair supervision. The finding that source-likelihood advantage does not reliably select beneficial local corrections raises the question of what supervision signal would, and whether task loss should be optimized directly rather than a logit-residual proxy.
  • Joint execution effects. Fitted corrections assume fixed parent inputs, but updates shift subsequent inputs; the reported input-response/direct-correction RMS ratios of 0.38–0.86 suggest a need for methods that account for these feedback effects.

Target Audience

Researchers and engineers working on model merging, sparse MoE architectures, router design and repair, and evaluation methodology for LLMs. It is also relevant to practitioners who need to decide whether a merged MoE checkpoint is safe to deploy, and to anyone designing diagnostics that distinguish structural change from functional harm.

Authors’ abstract

Model merging efficiently combines specialized large language models (LLMs) without joint retraining, but can substantially alter expert routing in Mixture-of-Experts (MoE) models. Such \emph{routing drift} is often interpreted as routing failure, raising a fundamental question that remains unclear: \emph{does routing drift after MoE merging actually indicate routing failure, and what evidence should justify repair?} We investigate these questions across DeepSeekMoE, OLMoE, and Qwen3-MoE proposing a routing analysis toolkit for controlled counterfactual interventions and token-level analysis. By crossing source and merged router inputs and parameters, we attribute most expert reassignments to input shifts rather than parameter changes at the same layer. However, source-relative routing differences poorly predict next-token likelihood gains from source-route restoration, and different expert selections can produce directionally similar mixture outputs. We therefore operationalize routing failure as \textit{task loss recoverable under a specified routing intervention, with non-routing parameters fixed.} These tests detect recoverable loss under deliberate router corruption, whereas source-route restoration does not establish reliable task benefits in the evaluated merged models. Motivated by these, we propose \emph{Selective Router Repair (SRR)} as a case study, and find that source-specialist token-likelihood advantages do not reliably identify beneficial local corrections. Together, these findings show that \textbf{routing drift alone is insufficient evidence of routing failure}: source-informed corrections must be judged by their task-level intervention effects. The analysis toolkit and SRR code are released.

Read the original paper