Skip to content
AI.info

Research

MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding

Overview Research area: On-device / edge multimodal machine learning systems, spanning computer vision, audio processing, and embedded sensing systems. Technical level: Intermediate. The paper is a sy

MMEdge: Accelerating On-device Multimodal Inference via Pipelined Sensing and Encoding
arXiv
2510.25327
Published
2025-10-29
Authors
Runxi Huang, Mingxuan Yu, Mingyu Tsoi, Xiaomin Ouyang

AI summary

Overview

  • Research area: On-device / edge multimodal machine learning systems, spanning computer vision, audio processing, and embedded sensing systems.
  • Technical level: Intermediate. The paper is a systems-and-efficiency paper; readers need some familiarity with neural encoders, sensor sampling rates, and latency budgets, but the core ideas are explained in accessible terms.
  • Scope: This paper proposes MMEdge, a system that pipelines sensor data acquisition together with feature encoding on resource-constrained devices, and adds runtime configuration optimization and cross-modal skipping to keep latency low without sacrificing accuracy.

What This Paper Is About

Conventional on-device multimodal systems wait until all sensor data in a time window (for example, one second of video and audio) is fully collected before running any model, which wastes time and forces faster modalities to idle while slower ones catch up. MMEdge instead chops the input into the smallest meaningful pieces (a video frame, an audio chunk), encodes each piece the moment it arrives, and stitches the results back together with a lightweight temporal aggregation module. The goal is real-time, privacy-preserving multimodal inference on edge hardware under strict latency constraints.

Key Contributions

  1. Motivation analysis: The authors measure where end-to-end latency actually goes in on-device multimodal pipelines and show that decomposing inference into fine-grained units is a viable path to acceleration, using an audio-visual speech recognition study on the Lip Reading in the Wild dataset.
  2. MMEdge framework: A pipelined sensing-and-encoding framework that processes each sensing unit immediately on arrival, paired with a lightweight temporal aggregation module that uses alternating temporal shift and multi-scale temporal difference features to recover temporal context.
  3. Adaptive multimodal configuration optimizer: An offline-profiled, online-searched optimizer that jointly selects sensing granularity (e.g., frame rate, audio chunk size) and model configuration (e.g., encoder size) per modality, using a lightweight accuracy predictor built on modality consistency and complementarity metrics under a latency constraint.
  4. Cross-modal speculative skipping: A mechanism that bypasses future units of slower modalities when an early prediction from faster modalities reaches sufficient confidence, reducing waiting time.
  5. Deployment: Evaluation on two public multimodal datasets and a real-world unmanned aerial vehicle (UAV) testbed for real-time human tracking, reporting a 75.83% reduction in end-to-end latency.

Main Findings

  • Idle waiting dominates the traditional pipeline: In the motivation study on an NVIDIA Jetson Xavier NX (16GB memory, 2-core, 10W power mode), video processing takes significantly longer than audio for both data collection and encoding, producing roughly 100 ms of idle waiting time before fusion.
  • Delays accumulate across samples: A 90 ms delay in the first sample postpones all subsequent samples, so latency grows cumulatively over time under the sequential framework.
  • Traditional vs. pipelined trade-off is stark: The traditional framework measured 242 ms latency at 92.76% accuracy; the naively pipelined framework measured 164 ms at 72.44% accuracy, reducing latency by roughly 80 ms but losing about 20% accuracy.
  • Sensing and model configurations are interdependent: Profiling 81 multimodal configurations (ResNet-18/34/50 video models; 20/25/29 FPS corresponding to 50/40/33 ms units; Small/Medium/Large audio models; 50/62.5/75 ms audio chunks) showed that increasing model depth and frame rate does not yield proportional accuracy gains. ResNet-34 at 25 FPS was comparable in accuracy to ResNet-50 at 20 FPS but with significantly lower latency; at a target of 85% accuracy, ResNet-34 at 25 FPS delivered high accuracy with minimal latency.
  • Cross-modal dependencies are real: Scaling both modalities together improved performance more than scaling one alone. Replacing ResNet-18 with ResNet-34 while reducing the audio model from medium to small improved both accuracy and latency.
  • Net deployment result: On the UAV testbed, MMEdge reduced end-to-end latency by 75.83% without compromising task performance.
  • Accuracy on the evaluation datasets: The paper states that MMEdge maintains high task accuracy across various system and data dynamics on two public multimodal datasets deployed on Nvidia edge devices. The specific dataset names and per-dataset accuracy figures are not reported in the available content (the text is truncated), and no accuracy number for the full MMEdge system on those datasets is given in the excerpt.

Methodology in Plain English

The traditional pipeline computes y = F(Σx): one large encoder consumes the entire temporal window. MMEdge reformulates this as y = Σf(x), where a lightweight encoder f runs on each small unit as it arrives. In video, that means 2D convolutions applied frame-by-frame instead of 3D convolutions over sequences; in audio, smaller models on short chunks. Because encoding happens during the sensing interval rather than after it, sensing and computation run concurrently and memory for buffering a whole window is avoided.

Breaking the input into units destroys temporal context, so the authors add two cheap fixes that introduce no extra latency. First, an alternating temporal shift: feature channels for unit i are split into groups (for example, 3), and the outer groups are replaced with features from neighboring units, X_i = [X_{i-k}^{(1)}, X_i^{(2)}, X_{i+k}^{(3)}], so each unit sees some past and future context. Boundary units keep their original values. Second, temporal difference features such as X_t − X_{t−1} and X_t − X_{t−2} are computed to capture short- and longer-term changes, then passed through a temporal encoder with global pooling.

For runtime adaptation, the system works in two stages. Offline, it profiles sensing and encoding latency across all candidate configurations under full end-to-end execution (so CPU scheduling and thermal throttling are reflected) and trains an accuracy predictor. Online, a greedy search picks the configuration maximizing predicted accuracy subject to a latency ceiling T_max, using a binary selection variable over sensing configuration c_s and model configuration c_m for each modality. Per-modality latency is modeled as L = max[L_E, L_S] × N + L_A (encoding vs. sensing interval, times the number of units, plus aggregation), and end-to-end latency as the slowest modality plus fusion.

Why This Matters

  • Research impact: The paper argues that sensing and inference cannot be optimized in isolation on a shared device, and it provides a concrete alternative framing — pipelining — plus evidence that joint cross-modal configuration search outperforms single-modality tuning.
  • Real-world applications:
    • Autonomous driving, where the paper notes perception and planning often must complete within 100 ms for timely control decisions.
    • Human-computer interaction, including audio-visual speech recognition and gesture recognition from streaming sensors.
    • Mobile health and activity monitoring, such as fall detection and systems like ADMarker that fuse depth, radar, and audio.
    • UAV-based human tracking, the actual deployment target demonstrated in this work.
  • Industry relevance: The system targets commodity edge hardware (NVIDIA Jetson Xavier NX class devices) and keeps all sensor data local, which matters for privacy-sensitive deployments and for scenarios with unpredictable network connectivity or bandwidth-heavy sensors such as LiDAR. Code is released at https://github.com/HKUST-MINSys-Lab/MMEdge.

Future Directions

  • Beyond two modalities: The optimization formulation is written for a general set of modalities M, but the reported experiments use audio-visual tasks. How well the pipelining and skipping mechanisms scale to three or more heterogeneous sensors is an open question.
  • Closing the temporal-accuracy gap: Naive pipelining cost about 20% accuracy in the motivation study. The temporal aggregation module is designed to recover this, but the available content does not report a per-dataset accuracy figure for the full system, so the exact residual gap is unclear.
  • Improving the accuracy predictor: The predictor must estimate accuracy for unseen samples and is trained offline. Its robustness to distribution shift, and how much its errors cost the optimizer, are natural next questions.
  • Generalizing beyond the profiled setup: Because latency profiles are device-specific and offline, re-profiling cost, portability to new hardware, and behavior under more severe thermal or contention dynamics remain open.

Target Audience

Researchers and practitioners working on edge AI, embedded sensing systems, and efficient multimodal machine learning; engineers deploying real-time perception or interaction models on resource-constrained devices; and graduate students interested in system-level co-design of sensing and inference. The paper is most useful to readers who already understand standard encoder architectures and latency budgeting, and who want a systems perspective on where time is actually spent in an on-device multimodal pipeline.

Authors’ abstract

Real-time multimodal inference on resource-constrained edge devices is essential for applications such as autonomous driving, human-computer interaction, and mobile health. However, prior work often overlooks the tight coupling between sensing dynamics and model execution, as well as the complex inter-modality dependencies. In this paper, we propose MMEdge, a new on-device multimodal inference framework based on pipelined sensing and encoding. Instead of waiting for complete sensor inputs, MMEdge decomposes the entire inference process into a sequence of fine-grained sensing and encoding units, allowing computation to proceed incrementally as data arrive. MMEdge also introduces a lightweight but effective temporal aggregation module that captures rich temporal dynamics across different pipelined units to maintain accuracy performance. Such pipelined design also opens up opportunities for fine-grained cross-modal optimization and early decision-making during inference. To further enhance system performance under resource variability and input data complexity, MMEdge incorporates an adaptive multimodal configuration optimizer that dynamically selects optimal sensing and model configurations for each modality under latency constraints, and a cross-modal speculative skipping mechanism that bypasses future units of slower modalities when early predictions reach sufficient confidence. We evaluate MMEdge using two public multimodal datasets and deploy it on a real-world unmanned aerial vehicle (UAV)-based multimodal testbed. The results show that MMEdge significantly reduces end-to-end latency while maintaining high task accuracy across various system and data dynamics. A video demonstration of MMEdge's performance in real world is available at https://youtu.be/qRew7sT-iWw.

Read the original paper