Skip to content
AI.info

Research

Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?

Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required? Overview Research area: Computer vision architectures for multi-channel imaging (vision transformers, mixtur

Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?
arXiv
2511.17400
Published
2025-11-21
Authors
Sukwon Yun, Heming Yao, Burkhard Hoeckendorf, David Richmond, Aviv Regev, Russell Littman

AI summary

Sparse Mixture-of-Experts for Multi-Channel Imaging: Are All Channel Interactions Required?

Overview

Research area: Computer vision architectures for multi-channel imaging (vision transformers, mixture-of-experts, efficient attention).

Technical level: Intermediate — the paper assumes familiarity with Vision Transformers, attention complexity, and the standard sparse mixture-of-experts formulation, though the core idea is explained clearly enough for a motivated non-specialist.

Scope: A proof-of-concept architecture, MoE-ViT, that treats each imaging channel as an expert and routes each patch to only a few channels, evaluated on two multi-channel datasets (JUMP-CP and So2Sat) for classification accuracy versus attention FLOPs. Published as a workshop paper at the AI for Science workshop (NeurIPS 2025), arXiv:2511.17400v1.

What This Paper Is About

Vision Transformers handle non-RGB, multi-channel data (like fluorescence microscopy or satellite imagery) by tokenizing each channel separately, which improves accuracy but makes attention cost grow quadratically in both the number of patches and the number of channels. The authors ask whether every channel really needs to interact with every patch, and propose a sparse mixture-of-experts design where a lightweight router selects only the most relevant channels per patch.

Key Contributions

  1. Reframes the problem from efficacy to efficiency. The paper identifies the overlooked computational bottleneck in cross-channel attention — cost growing from O(N²) to O(N²C²) — and poses the question of whether all channel interactions are necessary.
  2. Proposes MoE-ViT, an architecture that maps each channel to a distinct expert and adds a single-layer feed-forward channel router that selects the top-k channels per patch, reducing attention complexity to O(N²Ck) with k ≪ C (often 1–2).
  3. Introduces patch–channel cross-attention with a shared query projection and channel-specific key/value projections, where patches routed to a channel act as queries and patches within that channel act as keys and values.
  4. Demonstrates the approach on two real-world multi-channel domains (microscopy and satellite imagery), showing competitive or better accuracy at substantially lower attention FLOPs, with Top-k providing a tunable efficiency–accuracy knob.

Main Findings

  • Vanilla ViT is cheapest but least accurate. On JUMP-CP it reaches 56.24% top-1 accuracy at 0.29G attention GFLOPs; on So2Sat it reaches 58.93% at 0.02G.
  • Channel-wise baselines gain accuracy at high cost. ChAda-ViT reaches 68.16% on JUMP-CP and 63.94% on So2Sat; DiChaViT reaches 68.49% and 63.80% respectively. Both use 5.65G GFLOPs on JUMP-CP and 0.47G on So2Sat. The paper reports gains of +12.25% on JUMP-CP and +5.01% on So2Sat over the vanilla ViT.
  • Top-k = 2 balances the trade-off. On JUMP-CP, MoE-ViT (Top-k = 2) achieves 66.44% at 2.81G GFLOPs — roughly 50% less attention cost than ChAda-ViT for a 2.05% accuracy sacrifice. On So2Sat it achieves 64.66% at 0.36G GFLOPs, which the paper describes as a +0.72% accuracy improvement using 23% fewer GFLOPs.
  • Full-channel routing (Top-k = C) improves accuracy at matched cost. On JUMP-CP, MoE-ViT (Top-k = C) reaches 70.16% at 5.65G GFLOPs — matching the channel-wise baselines' attention cost while gaining +1.67% over the best baseline. The paper states a +0.72% gain on So2Sat as well, though the So2Sat table lists only Top-k = 1 and Top-k = 2 configurations.
  • Extreme sparsity (Top-k = 1) remains viable. MoE-ViT (Top-k = 1) reaches 64.06% at 2.33G GFLOPs on JUMP-CP and 63.12% at 0.35G on So2Sat, still beating the vanilla ViT.
  • Vanilla ViT activates more parameters than the sparsest MoE-ViT. On JUMP-CP, ViT uses 22.28M activated parameters versus 21.62M for MoE-ViT (Top-k = 1), because ViT concatenates all channels into a single higher-dimensional vector. MoE-ViT (Top-k = 2) uses 25.18M, and MoE-ViT (Top-k = C) uses 46.47M without a quadratic FLOP increase. On So2Sat the corresponding figures are 21.76M, 21.43M, and 24.98M.
  • Efficiency gains scale with spatial resolution. The advantages are more pronounced on JUMP-CP (224×224 images) than on So2Sat (32×32 crops).
  • Patch size dominates compute cost. In ablation, dropping patch size from 16 to 8 at Top-k = 2 increases attention GFLOPs from 2.81 to 22.61 on JUMP-CP (8×) and from 0.085 to 0.359 on So2Sat (4×). At fixed patch size on JUMP-CP, increasing Top-k from 1 to 4 raises GFLOPs from 15.02 to 22.61, 30.19, and 37.77, all still below the ChAda-ViT baseline, which exceeds 65 GFLOPs.

Methodology in Plain English

The authors start from the standard sparse mixture-of-experts idea: a router looks at an input and activates only the top-k most relevant experts, so you compute less while keeping a large pool of specialized capacity. They transplant that idea onto the channel dimension of a multi-channel image.

Concretely, an image with C channels is split into patches, and each patch is turned into a separate token per channel, with positional embeddings for spatial location and channel embeddings for channel identity. A tiny one-layer network (the channel router) scores all C channels for each patch token, and only the top-k scoring channels survive. The number of channel experts is set equal to the number of channels, so each expert specializes in one channel.

Then, for each surviving channel, the patches routed to it become the query matrix, while all patches belonging to that channel become the key and value matrices. Attention is computed between them, and each patch's output is aggregated from the experts it visited, weighted by router scores. Because the key and value projections are channel-specific, they play the role of the experts in the standard MoE formulation. The whole module replaces the standard multi-head attention in the Transformer encoder, leaving the rest of the ViT-Small backbone untouched.

The complexity analysis shows the attention term drops from O(BN²C²D) to O(BN²CkD), a k/C fraction of full attention, while the projection term drops from 3×O(BNCD²) to O(BNkD²) per projection type. Since attention scales with N² and projections with N, the attention savings dominate for N ≫ D.

Experiments follow the Channel-ViT setup: 100 epochs, ViT-Small backbones, AdamW, hierarchical channel sampling (HCS) during training, patch size 16 for JUMP-CP and 8 for So2Sat, reported as top-1 accuracy from the CLS token, with GFLOPs as a proxy for attention cost. Channel-ViT results are unavailable due to licensing restrictions, and ChAda-ViT was trained from scratch in a supervised setting without self-supervised pretraining.

Datasets: JUMP-CP is a microscopy dataset of 224×224 cell crops under chemical perturbations (plate BR00116991), with 5 fluorescence and 3 bright field channels (8 total), 127k training, 45k validation, and 45k test images, and a 161-class treatment group classification task. So2Sat contains satellite imagery from two remote sensors — 8 channels from Sentinel-1 and 10 from Sentinel-2 (18 total) — with 32×32 crops, 352k training and 24k test images, and a 17-class climate zone task.

Why This Matters

The paper opens a practical path for scaling vision transformers into domains where channel counts are large and growing, without paying quadratic attention costs. It shows that sparsity in the channel dimension is a controllable dial: Top-k can be tuned per compute budget, and larger k values can add model capacity without a quadratic FLOP penalty. Notably, larger Top-k activates more parameters at inference (46.47M at Top-k = C on JUMP-CP) while keeping complexity bounded, which the authors call arguably the most compelling aspect of the design.

Real-world applications:

  • Cell painting and fluorescence microscopy, where each stain or wavelength reveals structures invisible in other channels (JUMP-CP is the motivating example).
  • Satellite and remote sensing, where sensors such as Sentinel-1 and Sentinel-2 contribute complementary bands.
  • Digital pathology on large 2D multi-channel slides, listed as a promising extension.
  • 3D biological imaging datasets, also named as a target for extension.

Industry relevance: The authors are affiliated with the University of North Carolina at Chapel Hill and Genentech (gRED and the BRAID group), and the acknowledgments note Roche employment and equity for several authors. The efficiency framing targets large-scale and resource-constrained deployments where FLOPs and memory currently limit real-world use. The paper also points to production MoE libraries (DeepSeek-MoE, FastMoE) as a route to turning FLOP savings into wall-clock savings.

Future Directions

  1. Hardware-aware optimization. FLOPs are only a proxy; tighter architecture–hardware co-design and optimized MoE kernels could translate theoretical savings into real training and inference speedups.
  2. Increasing sparsity in the target matrix. A patch-specific router that allocates experts across spatial regions, possibly combined with batch-prioritized routing in the style of Riquelme et al., could sparsify further.
  3. Expanding to more biological and 3D datasets, including large 2D digital pathology and 3D imaging beyond the two datasets tested.
  4. Understanding the accuracy–routing trade-off. Open questions remain about why full-channel MoE routing outperforms the cross-channel baselines, and how routing behaves under different channel counts and spatial resolutions.

Target Audience

Researchers and engineers working on vision foundation models for scientific or remote-sensing imagery, especially those adapting ViTs to non-RGB, high-channel data. It is also useful for practitioners interested in mixture-of-experts efficiency techniques beyond language models, and for applied scientists in microscopy, pathology, or satellite analysis who need accuracy without prohibitive compute. Readers should have some background in transformer attention and its computational complexity to get the most out of it.

Authors’ abstract

Vision Transformers ($\text{ViTs}$) have become the backbone of vision foundation models, yet their optimization for multi-channel domains - such as cell painting or satellite imagery - remains underexplored. A key challenge in these domains is capturing interactions between channels, as each channel carries different information. While existing works have shown efficacy by treating each channel independently during tokenization, this approach naturally introduces a major computational bottleneck in the attention block - channel-wise comparisons leads to a quadratic growth in attention, resulting in excessive $\text{FLOPs}$ and high training cost. In this work, we shift focus from efficacy to the overlooked efficiency challenge in cross-channel attention and ask: "Is it necessary to model all channel interactions?". Inspired by the philosophy of Sparse Mixture-of-Experts ($\text{MoE}$), we propose MoE-ViT, a Mixture-of-Experts architecture for multi-channel images in $\text{ViTs}$, which treats each channel as an expert and employs a lightweight router to select only the most relevant experts per patch for attention. Proof-of-concept experiments on real-world datasets - JUMP-CP and So2Sat - demonstrate that $\text{MoE-ViT}$ achieves substantial efficiency gains without sacrificing, and in some cases enhancing, performance, making it a practical and attractive backbone for multi-channel imaging.

Read the original paper