Research
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline Analysis Overview Research area: Efficient machine learning / on-device (edge) inference of small language models, combining per

- arXiv
- 2602.11506
- Published
- 2026-02-12
- Authors
- Zhen Bi, Xueshu Chen, Luoyang Sun, Yuhang Yao, Qing Shen, Jungang Lou, Cheng Deng
AI summary
RooflineBench: A Benchmarking Framework for On-Device LLMs via Roofline AnalysisOverview
Research area: Efficient machine learning / on-device (edge) inference of small language models, combining performance modeling with empirical benchmarking.
Technical level: Intermediate. The paper assumes familiarity with Transformer inference (prefill vs. decoding, KV cache, attention variants) and basic hardware concepts (FLOPS, memory bandwidth), but its central tool — the Roofline model — is explained from first principles.
Scope (one sentence): The paper proposes a Roofline-based benchmarking framework that measures how close on-device LLM decoding runs to the physical limits of heterogeneous edge hardware, and introduces the Relative Inference Potential metric to compare architectures on the same device.
What This Paper Is About
Deploying language models on phones, laptops, and single-board computers is bottlenecked by hardware physics rather than by model capability, yet existing benchmarks mostly report raw latency or throughput without saying how far those numbers are from what the hardware could theoretically achieve. The authors build a runtime-integrated profiling framework that measures each device's real peak compute and memory bandwidth, then places observed LLM inference points on a Roofline plot to expose exactly where the inefficiency lies. Their goal is to give hardware and model designers a shared, quantitative language for on-device efficiency, and to show how sequence length, model depth, precision, and attention design each move a model closer to or further from the hardware ceiling.
Key Contributions
- An integrated Roofline-based benchmarking framework that unifies architectural primitives (linear, attention, and FFN layers) with hardware constraints through Operational Intensity (OI). It defines an "inference potential region" and introduces the Relative Inference Potential (Φ) as a new metric for comparing two LLMs on the same hardware substrate.
- A comprehensive empirical characterization across heterogeneous compute tiers, showing that inference efficiency is primarily governed by context length and attention architecture, while identifying a critical OI regression at increased model depth.
- Hardware-software co-design guidance, including identification of an "efficiency trap" caused by hardware heterogeneity and a demonstration that structural refinements such as Multi-head Latent Attention (MLA) unlock latent inference potential across different substrates.
- An open release: the code is available at https://github.com/banbu-ai/roofline_bench (stated in Appendix C).
Main Findings
- Decoding is memory-bandwidth bound. For a decoder-only model, the forward pass costs roughly 2·n_params FLOPs per token, giving a decoding Operational Intensity of approximately 2/b_prec — about 1 FLOP/Byte at 16-bit precision. Because edge accelerators have peak-compute-to-bandwidth ratios far exceeding 1, computational cores sit largely idle waiting for tensors.
- Context length dominates efficiency. Across all tested models, the LISO (Long In, Short Out) scenario consistently achieves the highest execution efficiency and sits closest to the Roofline ridge point, while SILO (Short In, Long Out) resides deep in the memory-bound regime with severe hardware underutilization. LISO amortizes fixed weight-loading overhead and approaches the compute-bound limit.
- Operational intensity peaks at shallow depth. Scaling Transformer layers from 2 to 64 produces a non-monotonic "arch" trajectory in OI, with a clear inflection at 3–5 layers. Beyond that, cumulative memory bandwidth pressure from weight streaming outpaces marginal gains in arithmetic reuse, so OI retreats and the model is forced deeper into the memory-bound regime.
- Quantization helps most where memory is the bottleneck. Moving from FP16 to Q8_0 to Q4_K_M shifts data points diagonally toward the upper-right of the Roofline plane. The gain is dramatic in memory-bound SILO but much smaller in LISO (top-right), where clusters for the three precisions appear noticeably tighter and execution is limited by peak compute rather than data movement.
- MLA outperforms MHA and GQA at matched scale. With all LLMs scaled to approximately 1.5B parameters, MLA consistently achieves the highest OI and attainable GFLOPS across all four sequence scenarios; at this scale on edge hardware, GQA exhibits the lowest efficiency of the three. Latent compression of the KV cache reduces per-step data movement, shifting execution closer to the Roofline ridge.
- Hardware ridge points vary enormously. The theoretical ridge points of five representative devices range from 8.98 to 38.00 FLOPs/Byte: RTX 3090 at 38.00, RTX 3070Ti Laptop at 37.05, Apple M1 Pro at 25.39, Jetson Orin Nano at 20.48, and Raspberry Pi 5 at 8.98. Theoretical peak performance spans three orders of magnitude, from 10^11 to 10^13 FLOPs/sec.
- The "efficiency trap." Because ridge points differ so much across devices, a single architecture (e.g., a 1.5B model in LISO) may reach near-optimal saturation on a Jetson Orin or Raspberry Pi 5 yet remain severely underutilized on an RTX 3090.
- Architectural gains are preserved across platforms. When scaling from Apple M1 Pro to RTX 3070Ti Laptop in the compute-intensive LISO scenario, MLA, MHA, and GQA all shift upward and rightward in parallel, but MLA maintains the superior baseline position — suggesting MLA's latent compression is a robust efficiency floor rather than a platform-specific trick.
- Measured versus theoretical peaks can diverge. The paper's Table 1 reports, for example, RTX 3090 bandwidth of 936.20 / 560.02 GB/s (theoretical / measured), FP16 peak of 35.58 / 66.20 TFLOPS, and FP32 peak of 35.58 / 24.28 TFLOPS. The authors note that measured FP16 peak performance for NVIDIA RTX 30 series GPUs and Jetson Orin Nano significantly exceeds theoretical FP32 peaks because specialized hardware such as Tensor Cores activates during evaluation.
Methodology in Plain English
The authors avoid simulation. Instead, they build an instrumented runtime that watches real inference and compares it against limits they measure on the actual device.
The measurement proceeds in three parts:
- Count the work analytically, not with hardware counters. Directly reading FLOPs from hardware counters is imprecise due to instrumentation overhead, so the authors use an analytical formulation to estimate theoretical FLOPs for heterogeneous architectures such as MHA and GQA. Given a hidden dimension H, sequence length N, and head configuration {n_q, n_k, n_v}, they compute the computational load for the Linear, Attention, and FFN layers per decoding step, ensuring a uniform standard across backends.
- Estimate data movement. Total memory traffic Q per generated token is approximated as the summation of the model parameter size and the active KV cache entries read/written in that step — an approach intended to hold whether the device uses Unified Memory (Apple Silicon) or dedicated VRAM (CUDA devices). Performance is W/T (FLOPs over measured end-to-end latency) and OI is W/Q.
- Measure the hardware ceiling empirically. They profile achievable peak memory bandwidth (BW_peak) and compute performance (P_peak) directly: Unified Memory bandwidth for Apple Silicon, and dedicated Video Memory bandwidth for CUDA-enabled devices such as the RTX 4090, since that is the primary bottleneck for on-device inference.
The standard Roofline bound is P = min(P_peak, OI × BW_peak). Because autoregressive decoding is inherently memory-bound, the analysis concentrates on the sloped region of the graph.
The Relative Inference Potential (Φ). This metric is the distance from an observed point P(OI_p, Perf_p) to the hardware ridge R(OI_r, π), and its formula changes depending on which side of the ridge the point sits:
- In the memory-bound regime (OI < OI_r), Φ is the Euclidean distance: the square root of (OI_r − OI_p)² + (π − Perf_p)², capturing the dual need to raise both operational intensity and throughput.
- In the compute-bound regime (OI ≥ OI_r), horizontal OI gains barely raise the ceiling, so Φ is the vertical distance π − Perf_p.
- Points on opposite sides of the ridge are declared fundamentally incomparable, because they face distinct physical bottlenecks, necessitating regime-specific evaluation.
What was tested. The evaluation uses four sequence patterns reflecting edge use cases: SISO (Short In, Short Out), SILO (Short In, Long Out), LISO (Long In, Short Out), and LILO (Long In, Long Out). Edge-scale models for the Raspberry Pi 5 tier include Pythia (160M, 410M), Qwen2.5-0.5B, and SmolLM2 (135M, 360M), primarily under Q8_0 quantization. Mobile-class models (0.6B to 1.8B) span attention types: GQA for Qwen2.5-1.5B, Llama-3.2-1B, Qwen3-0.6B, and Fox-1-1.6B; MHA for SmolLM2-1.7B; and MLA for PLM-1.8B. Four hardware platforms are evaluated directly — Apple M1 Pro, RTX 3070 Ti Laptop, Jetson Orin Nano Super, and Raspberry Pi 5 — with Figure 2 additionally covering Qwen2.5-Instruct (GQA), Llama-3.2-Instruct (GQA), PLM-Instruct (MLA), and SmolLM2-Instruct (MHA) in FP16 on Apple M1 Pro. No dataset sizes or accuracy benchmarks are reported; the paper measures inference efficiency, not task quality.
Why This Matters
Most on-device LLM benchmarks report wall-clock numbers in isolation. RooflineBench instead contextualizes those numbers against each device's measured physical ceiling, so a throughput figure on a Raspberry Pi 5 and one on an RTX 3090 can be compared in terms of fraction of potential achieved rather than raw speed. This separates software-level optimization from inherent hardware capability and gives architectural claims (e.g., "MLA is better") a physical, falsifiable basis.
Real-world applications, drawn from the paper's own scenario taxonomy:
- SISO (Short In, Short Out): local voice commands and similar latency-critical interactions.
- SILO (Short In, Long Out): creative writing and code completion, where the paper finds the most acute memory-bound underutilization.
- LISO (Long In, Short Out): RAG-based context extraction and document summarization — the scenario that comes closest to saturating compute.
- LILO (Long In, Long Out): document translation, where both context and generation are long.
Industry relevance: The findings feed directly into decisions about quantization level selection, how many layers a mobile model should have, which attention mechanism to ship, and which silicon to target. The paper argues that hardware heterogeneity means general-purpose logic frequently fails to exploit modern architecture structures, so specialized silicon support for critical primitives such as MLA and sparse attention is needed to let a model's theoretical OI reach its full potential. It also notes that MLA and sparse attention appear in shipping systems such as DeepSeek V2, PLM, and MiniCPM, where MiniCPM's sparse attention processes less than 5% of tokens in long-context scenarios.
Future Directions
- Extend to Mixture of Experts architectures. Quantifying theoretical FLOPs for MoE during dynamic inference remains difficult because stochastic token routing complicates estimation of the execution ceiling. Extending the Roofline model to sparse activation regimes is flagged as essential for finding the efficiency traps unique to localized MoE deployment.
- Broaden attention mechanism coverage. The authors plan to adapt the methodology to emerging structural variants and hybrid architectures beyond the MLA and GQA patterns analyzed here.
- Widen the hardware scope. They intend to test a more diverse array of heterogeneous edge devices to further validate cross-platform robustness.
- Study software stack variance. Different inference engines — TensorRT LLM, vLLM, and ONNX Runtime are named — introduce substantial performance variance. Investigating how these stacks influence OI and realized throughput is framed as necessary for a holistic view of the deployment pipeline and for true hardware-software co-design.
Target Audience
This paper is most useful to (a) edge/on-device ML engineers deciding how to quantize, size, and architect models for specific hardware; (b) hardware architects and accelerator designers looking for evidence about which primitives (latent compression, sparse attention) deserve dedicated silicon; (c) systems researchers working on benchmarking methodology who want a physics-grounded alternative to raw throughput reporting; and (d) graduate students or practitioners already comfortable with Transformer inference who want a worked example of applying the Roofline model to language model decoding. Readers expecting accuracy evaluations or dataset-level results will not find them here — the framework is about execution efficiency against hardware limits, not model intelligence.
Authors’ abstract
The transition toward localized intelligence through Small Language Models (SLMs) has intensified the need for rigorous performance characterization on resource-constrained edge hardware. However, objectively measuring the theoretical performance ceilings of diverse architectures across heterogeneous platforms remains a formidable challenge. In this work, we propose a systematic framework based on the Roofline model that unifies architectural primitives and hardware constraints through the lens of operational intensity (OI). By defining an inference-potential region, we introduce the Relative Inference Potential as a novel metric to compare efficiency differences between Large Language Models (LLMs) on the same hardware substrate. Extensive empirical analysis across diverse compute tiers reveals that variations in performance and OI are significantly influenced by sequence length. We further identify a critical regression in OI as model depth increases. Additionally, our findings highlight an efficiency trap induced by hardware heterogeneity and demonstrate how structural refinements, such as Multi-head Latent Attention (MLA), can effectively unlock latent inference potential across various hardware substrates. These insights provide actionable directions for hardware-software co-design to align neural structures with physical constraints in on-device intelligence. The released code is available in the Appendix C.