Research
A Study on Inference Latency for Vision Transformers on Mobile Devices
Overview Research area: On-device computer vision — specifically, inference latency of Vision Transformers (ViTs) on mobile hardware, compared against convolutional neural networks (CNNs). Technical l

- arXiv
- 2510.25166
- Published
- 2025-10-29
- Authors
- Zhuojin Li, Marco Paolieri, Leana Golubchik
AI summary
Overview
Research area: On-device computer vision — specifically, inference latency of Vision Transformers (ViTs) on mobile hardware, compared against convolutional neural networks (CNNs).
Technical level: Intermediate. The paper is accessible to readers familiar with basic neural network concepts (attention, convolution, quantization, FLOPs), and it presents empirical measurement results rather than new model architectures.
Scope (1 sentence): The paper measures, characterizes, and predicts the on-device inference latency of 190 real-world and 1000 synthetic ViTs across 2 machine learning frameworks and 6 mobile platforms, using the resulting dataset to train latency predictors accurate enough for practical use.
What This Paper Is About
Vision Transformers achieve strong accuracy on computer vision tasks, but their self-attention mechanism is computationally expensive and poorly suited to memory-constrained mobile devices. Prior work in this space studied only a small number of ViTs (for example, 9 ViTs in one earlier study) and relied on a single ML framework, leaving the practical behavior of modern efficient ViTs on phones largely unexamined. This paper addresses that gap by quantitatively measuring how ViT architecture choices, memory formats, activation functions, and ML framework implementations affect real inference latency on mobile devices, then building a dataset and predictors that let developers estimate latency of new ViTs without deploying them.
Key Contributions
-
A quantitative comparison of 190 real-world ViTs against 102 real-world CNNs on mobile platforms, covering differences in latency, performance bottlenecks (memory- versus compute-bound behavior), and memory consumption, along with explanations of the underlying causes (memory formats, activation function selection, and ML framework implementations).
-
A search space for generating synthetic ViTs built from representative building blocks used in state-of-the-art efficient ViTs, and a released dataset containing profiling information for 1000 synthetic ViTs and 190 real-world ViTs across 6 mobile platforms and 2 mainstream ML frameworks, covering different CPU core combinations and data representations.
-
Demonstration that simple ML latency predictors trained on 900 synthetic ViTs work in practice: error rates of 4.4% on PyTorch Mobile and 4.8% on TFLite for mobile CPUs on synthetic ViTs, and 8.2% on PyTorch Mobile and 6.1% on TFLite for real-world ViTs on mobile CPUs, supporting use cases such as Neural Architecture Search (NAS) and collaborative (split) inference.
-
Evidence that predictors trained only on synthetic ViTs transfer to unseen real-world ViT architectures, validating the representativeness of the synthetic search space.
Main Findings
-
ViTs are slower than CNNs at comparable FLOPs: Among models with fewer than 5 GFLOPs, a ViT with 5.00 GFLOPs incurs 1.75x the latency of a CNN with 4.95 GFLOPs. The cause is attributed to architectural differences: self-attention has complexity O(H²W²) for an image of dimensions H×W, versus O(HW) for convolutions.
-
Linear and activation operations dominate ViT latency: In CNNs most end-to-end latency comes from convolutions, whereas in ViTs a significant portion comes from linear operations (the core of self-attention blocks) and from activation operations, primarily GELU.
-
GELU latency depends on input values: The Gaussian error function
erfis computed differently depending on its absolute input values, with discontinuities corresponding to GELU input values {1.19, 1.77, 4.04, 5.66}. The latency for an input value of 2 is 2.85x longer than for an input value of 1. The approximate tanh-based GELU also depends on input values. -
ViTs are more memory-bound than CNNs: Arithmetic intensity of ViTs is generally lower than that of CNNs on Pixel 4. Raising memory frequencies from lowest to highest produces speedups over 3.40x for 75% of ViTs versus only 28% of CNNs; boosting CPU clock from 0.83 GHz to 2.84 GHz gives speedups under 1.93x for 89% of ViTs versus only 12% of CNNs.
-
Memory format significantly changes convolution latency: Channel-last (NHWC) input tensors produced average speedups of 2.21x and 1.58x over channel-first (NCHW) for tensor shapes of 64×64 and 48×48, respectively, on PyTorch Mobile.
-
ViTs consume more memory than comparable CNNs, and scale worse with resolution: Vanilla ViT-S (patch size 32) consumed 14% more memory than ResNet-18 (width scale 0.5) at 224×224 pixels, and the gap grew to 43% at 512×512 pixels. The object detection ViT DETR-ResNet101 consumes 5.0 GB of mobile memory for a single 512×1333 image, compared against 6 GB of RAM on an iPhone 15 and 80 GB of GPU memory on an Nvidia H100.
-
Framework implementation matters: Depthwise convolution latency in PyTorch Mobile increases non-linearly with spikes when input channels are a multiple of 32, and latency for 64×64 input shapes is surprisingly greater than for 72×72 shapes at the same channel count, attributed to memory format conversion between PyTorch Mobile and XNNPACK. DWConv in TFLite is significantly faster and scales linearly.
-
Quantization helps on efficient cores but not on powerful cores in PyTorch Mobile: After quantization, ViTs accelerate on small cores but show degradation on large and medium cores, partly because the quantized library QNNPACK is less efficient for linear operations than the floating-point XNNPACK library on powerful cores. Activation functions improve substantially in TFLite but degrade in PyTorch Mobile, and normalization degrades in both.
-
Non-linear predictors are accurate, linear ones are not: Random Forest and GBDT reach a maximum of 4.9% error on convolution and linear operations on synthetic ViTs, while Lasso reaches 17.9% error on convolution operations in PyTorch Mobile. GBDT reaches 2.1% error on PyTorch Mobile GPUs and, as reported in the results section, 8.6% on TFLite GPUs (the abstract and contributions list report 8.9% for TFLite GPUs), with Lasso performing worse on both.
-
Synthetic-trained predictors generalize to real-world ViTs: GBDT maximum MAPEs on real-world ViT multicore CPU predictions were 11.0% on Snapdragon 855, 13.3% on Exynos 9820, 13.7% on Snapdragon 710, 14.2% on Helio P35, 13.6% on A12 Bionic, and 14.9% on A14 Bionic. On synthetic ViTs the corresponding maximum errors were 8.2%, 12.5%, 6.9%, 9.1%, 10.6%, and 11.1%.
-
Errors rise with heterogeneous core usage: The worst predictions occur when many cores — especially heterogeneous combinations such as 1 large and 2 medium cores on Exynos 9820 — are used, attributed to resource contention, inter-cluster communication, and thread synchronization overhead.
-
Quantized integer representations predict better than floating point in some cases: Prediction errors were generally lower for integer representations after quantization than floating-point, for example on A12 Bionic in TFLite, because GELU mispredictions matter less when activations are a smaller share of end-to-end latency.
Methodology in Plain English
The researchers collected 190 real-world ViTs from Timm and HuggingFace's Transformers and converted them into PyTorch Mobile format, then deployed them on six phones: Google Pixel 4 (Snapdragon 855), Motorola One Fusion (Snapdragon 710), Samsung Galaxy S10 (Exynos 9820), Samsung Galaxy A03s (Helio P35), Apple iPhone 12 (A14 Bionic), and Apple iPhone XS (A12 Bionic). They measured latency across various CPU core combinations, with each core assigned a thread, and also evaluated quantization; 64 of the 190 ViTs could not be quantized due to unsupported operations in PyTorch Mobile, and only 25 of the 190 had TensorFlow implementations for TFLite comparison. All models were run on mobile CPUs for the main comparison because many ViT operations, such as the roll operation in Swin, are unavailable on mobile GPUs in current frameworks.
To understand why latency behaves as it does, they profiled per-operation latency and memory, used Performance Monitoring Unit counters on ARM processors to compute arithmetic intensity, and ran controlled experiments varying memory frequency through DVFS governors and CPU clock speed.
Because real-world ViTs are hard to measure on mobile GPUs, they designed a search space of synthetic ViTs built from representative blocks (SepConv or Attention token mixers, batchnorm or layernorm, GELU or SiLU, varying embedding lengths, MLP expansion ratios, filter shapes, attention heads, and spatial-reduction attention). From this space they sampled 1000 synthetic ViTs and measured them across 2 frameworks and 6 devices. They trained per-operation latency predictors using Lasso, Random Forest, and Gradient Boosted Decision Trees with MAPE as the loss, splitting the data into 900 synthetic ViTs for training and 100 for testing, then tested the same predictors on the 190 real-world ViTs without retraining.
Why This Matters
Impact on research: The study supplies a large, multi-device, multi-framework latency dataset and shows that FLOPs — a common proxy for cost in architecture design — is an unreliable latency indicator for ViTs, in part because GELU latency varies with input values. It also shows that CNN-derived latency intuition does not transfer cleanly to ViTs, which are more memory-bound.
Real-world applications:
- Neural Architecture Search: latency prediction avoids repeated deployment to phones and lets search enforce latency constraints on a target device.
- Collaborative (split) inference: predicting latency for each part of a model helps choose how to partition computation between the device and a cloud server, balancing saved local computation against transmission cost.
- On-device augmented reality and motion analysis: knowing which ViT designs fit within mobile memory and latency budgets guides model selection for real-time processing.
- Framework and kernel engineering: the memory-format and quantization findings point directly at specific conversion overheads and library choices that developers can act on.
Industry relevance: The findings give mobile ML engineers concrete guidance — prefer channel-last memory formats, be cautious about quantizing on powerful cores in PyTorch Mobile, and expect different latency behavior between PyTorch Mobile and TFLite for identical operations. Since most state-of-the-art ViTs are implemented in PyTorch, the emphasis on PyTorch Mobile, with TFLite as a comparison point, reflects actual deployment practice.
Future Directions
-
Extending GPU coverage: Measurements for synthetic ViTs on Mali G76 and Helio P35 could not be collected due to out-of-GPU-memory issues, so GPU analysis covers only two Android GPUs with PyTorch Mobile and four with TFLite. Restoring coverage would test whether the predictors hold across all six platforms on GPU.
-
Improving predictions for heterogeneous multicore execution: The largest errors occur with mixed core combinations, which the authors attribute to contention and synchronization overhead. Modeling these effects more directly is an open problem.
-
Making activation latency predictable: Because GELU latency depends on input values and current predictors use only operation configurations, activation errors remain comparatively high. Input-aware or value-sensitive modeling could close this gap.
-
Broadening the framework and model space: Only 25 of the 190 real-world ViTs had TensorFlow implementations, and 64 could not be quantized in PyTorch Mobile, so the datasets used for framework comparisons and quantization studies are partial subsets of the full collection.
Target Audience
Mobile and embedded machine learning engineers, computer vision researchers working on efficient architectures, and practitioners building systems that require on-device inference within strict memory and latency budgets. It is also relevant to researchers working on Neural Architecture Search and collaborative inference, who need accurate latency estimates as part of their optimization objectives, and to framework developers interested in how memory-format handling and quantized kernels affect real measured performance.
Authors’ abstract
Given the significant advances in machine learning techniques on mobile devices, particularly in the domain of computer vision, in this work we quantitatively study the performance characteristics of 190 real-world vision transformers (ViTs) on mobile devices. Through a comparison with 102 real-world convolutional neural networks (CNNs), we provide insights into the factors that influence the latency of ViT architectures on mobile devices. Based on these insights, we develop a dataset including measured latencies of 1000 synthetic ViTs with representative building blocks and state-of-the-art architectures from two machine learning frameworks and six mobile platforms. Using this dataset, we show that inference latency of new ViTs can be predicted with sufficient accuracy for real-world applications.