Research
From Loop Nests to Silicon: Mapping AI Workloads onto AMD NPUs with MLIR-AIR
Overview Research area: Compiler infrastructure for AI accelerators — specifically an MLIR-based intermediate representation (IR) for mapping machine-learning workloads onto spatial hardware, demonstr

- arXiv
- 2510.14871
- Published
- 2025-10-16
- Authors
- Erwei Wang, Samuel Bayliss, Andra Bisca, Zachary Blair, Sangeeta Chowdhary, Kristof Denolf, Jeff Fifield, Brandon Freiberger, Erika Hunhoff, Phil James-Roxby, Jack Lo, Joseph Melber, Stephen Neuendorffer, Eddie Richter, Andre Rosti, Javier Setoain, Gagandeep Singh, Endri Taka, Pranathi Vasireddy, Zhewen Yu, Niansong Zhang, Jinming Zhuang
AI summary
Overview
Research area: Compiler infrastructure for AI accelerators — specifically an MLIR-based intermediate representation (IR) for mapping machine-learning workloads onto spatial hardware, demonstrated on AMD's Neural Processing Units (NPUs).
Technical level: Advanced. The paper assumes familiarity with compiler IRs, MLIR dialects, dataflow architectures, DMA engines, and loop-nest transformations. The case studies (matrix multiplication, attention) are accessible, but the core material is aimed at compiler and accelerator architects.
Scope: The paper introduces MLIR-AIR, an open-source MLIR dialect and end-to-end compilation flow that turns high-level loop nests from AI frameworks into explicitly scheduled spatial programs for AMD NPUs, validated on matrix multiplication and a multi-head attention block from LLaMA 2.
(Note: the paper is listed on arXiv under cs.CL / Natural Language Processing, though its subject matter is compiler and hardware architecture.)
What This Paper Is About
General-purpose compilers hide parallelism, data locality, and synchronization behind abstractions like shared memory and thread scheduling, which limits how well they can exploit modern spatial accelerators made of many distributed compute tiles, partitioned memories, and message-passing interconnects. The authors build MLIR-AIR so that placement of computation, movement of data, and ordering of execution become explicit, compiler-visible constructs rather than being deferred to hardware schedulers or hand-written runtime coordination. The goal is to let a compiler take structured loop nests from high-level AI frameworks and lower them into statically scheduled, tiled spatial programs that overlap communication with computation on AMD NPUs.
Key Contributions
-
The AIR dialect itself. A new MLIR dialect that exposes spatial and temporal program structure. It models spatial partitioning with
air.herd, point-to-point communication withair.channel, and explicit synchronization withair.token, plus hierarchical scheduling constructs (air.launch,air.segment) and data-movement constructs (air.memcpy,air.channel.put/air.channel.get). The dialect lets the compiler coordinate computation, data movement, and synchronization without forcing users down to vendor-specific abstractions. -
A complete end-to-end compiler flow. MLIR-AIR lowers workloads written in high-level Python frameworks through MLIR's SCF and LinAlg dialects and the AIR dialect, then into low-level code for AMD NPUs via MLIR-AIE, dispatched using the NPU runtime. The stack is open source and modular, integrating with the wider MLIR ecosystem.
-
Demonstration on two representative AI workloads. Matrix multiplication and the multi-head attention block from LLaMA 2, showing that MLIR-AIR produces statically scheduled programs that exploit locality, parallelism, and pipelining on tiled hardware.
-
A vendor-agnostic design. Because AIR primitives target features common to a class of spatial accelerators, the framework is intended as a foundation for targeting accelerators beyond AMD's NPU, with frontends such as PyTorch, TensorFlow, and Triton supported through Torch-MLIR, TOSA, and Triton-Shared.
Main Findings
- Matrix multiplication efficiency: MLIR-AIR achieves up to 78.7% compute efficiency on matrix multiplication.
- Near hand-optimized performance: For matrix multiplication, MLIR-AIR generates implementations whose performance is almost identical to state-of-the-art, hand-optimized matrix multiplication written using the lower-level, close-to-metal MLIR-AIE framework.
- Compact fused attention: The AIR interface supports fused multi-head attention implementations using approximately 150 lines of code, which the authors present as evidence that complex workloads can be expressed tractably while still mapping efficiently to spatial hardware.
- Token-based dependency tracking: Fine-grained asynchronous parallelism is captured through an asynchronous control and dataflow graph (ACDG) embedded in the IR via production and consumption of
air.tokenvalues. Tokens are SSA values encoding read-after-write, write-after-read, and write-after-write dependencies, so correctness checks and transformation legality reuse MLIR's native verification infrastructure. - Decoupled data movement: Instead of a single coupled copy construct,
air.channel.putandair.channel.getseparate the two ends of a transfer, linked by a symbolicair.channel. This lets communication be scheduled alongside computation when useful, or independently when that is better for performance or modularity, and enforces backpressure-based synchronization. - Hardware trends that motivate the design: The paper identifies six trends in spatial hardware — complex system hierarchy, dispatch placement, multi-root memory hierarchy, peer memory movement, data movement offload, and asynchronous execution — and derives four corresponding compiler needs: exposing memory-allocation controls across hierarchy levels and NUMA domains, exposing compute-placement controls, separating data movement from computation in the IR, and enabling dependency resolution close to hardware.
- Deterministic target behavior: The AMD NPU has no caches for data or instructions and computes on local tiles of data, eliminating memory access latency variation. The paper argues this predictability is what lets tiling compilers construct efficient dataflow with high utilization.
- Resource-scoped scheduling:
air.herdoperations are scheduled atomically — only when resources for all workers are available — and must guarantee independent forward progress for their workers, with workers allocated as a physically local contiguous block.air.launchcan release unused resources back to the runtime once nested work begins executing. - Profiling support: MLIR-AIR emits execution traces capturing performance metrics during hardware execution, visualized with tools such as Chrome Tracing or Perfetto UI, so developers can inspect runtime parallelism between computation and data movement.
Methodology in Plain English
The authors start from an observation about hardware: modern accelerators are built from many compute tiles, several separate memory levels, and dedicated data-movement engines, and getting good performance on them depends on where work runs and where data lives. Thread-centric programming models deliberately hide those decisions, which is convenient but leaves performance on the table.
Their approach is to make those decisions explicit in a compiler intermediate representation. They define AIR as an MLIR dialect, so it composes with existing MLIR dialects rather than replacing them. Computation and kernel bodies continue to be described by standard MLIR constructs; AIR adds the spatial scaffolding around them.
The flow works in layers. High-level AI frameworks (PyTorch, TensorFlow, Triton) arrive through MLIR frontends (Torch-MLIR, TOSA, Triton-Shared). These lower into MLIR's structured control flow (SCF) and linear algebra (LinAlg) dialects to keep loop nests and tensor semantics intact. Then the AIR dialect takes over, describing asynchronous execution and memory-hierarchy interaction. From there, platform-specific backends — MLIR-AIE for NPUs, or LLVM-based pipelines for CPUs and GPUs — generate hardware-specific control and dataflow, deployed through runtimes such as XRT or ROCr.
Within AIR, the authors group operations into three families: scheduling constructs (air.launch, air.segment, air.herd) that distribute work hierarchically across compute resources; data locality constructs (air.memcpy, air.channel) that describe transfers aligned with memory hierarchy and DMA affinity; and synchronization constructs (air.token) that express operation-level and loop-carried dependencies. A key design choice is splitting scheduling responsibility between compiler and runtime: the compiler defines tightly coupled concurrent groups ("herds"), while the runtime has flexibility to place those groups on devices of varying size across hardware generations.
To evaluate, they compile matrix multiplication and the LLaMA 2 multi-head attention block, comparing the matmul result against hand-optimized MLIR-AIE code and reporting compute efficiency, and reporting the source-line count needed for the fused attention implementation.
Why This Matters
Impact on research. The paper argues that as architectures grow more spatial and asynchronous, the thread-centric abstraction becomes a bottleneck rather than a convenience — it requires dense interconnects, deep caches, and complex runtimes, all of which cost energy, silicon area, and design complexity. MLIR-AIR offers a concrete alternative: an IR where placement, data sharing, and ordering are programmable. It also positions itself against polyhedral compilers (which it extends beyond loop transformation into asynchronous scheduling and data movement) and against template-based or runtime-coordinated frameworks, by keeping spatial scheduling and asynchronous execution inside the IR. Because the dialect is vendor-agnostic and open source, it provides a reusable foundation for other spatial accelerators.
Real-world applications:
- Dense linear algebra on NPUs. Efficient matrix multiplication is the workhorse of essentially every deep-learning model, and the paper shows the flow can approach hand-tuned quality on this kernel.
- Transformer inference. The multi-head attention block from LLaMA 2 is the paper's second case study; fused attention implementations matter directly for large language model serving, particularly where attention dominates latency.
- Client and edge AI on NPUs. AMD's NPU line targets low-power, low-latency inference on laptops and embedded systems, where explicit data movement and no-cache determinism matter a great deal.
- Retargeting to non-AMD spatial hardware. The stacked architecture (frontends, AIR, pluggable backends, runtimes) is designed so that the same high-level programs can be lowered to other tiled accelerators.
Industry relevance. The work comes from AMD's Research and Advanced Development group as an open-source contribution (Xilinx/mlir-air), and its competitive baseline is AMD's own lower-level MLIR-AIE framework. That framing matters: the pitch is not that a high-level abstraction can beat hand-tuned code, but that it can come close while remaining portable and far more compact to write. For hardware vendors, that means a compiler stack that can adapt to new accelerator generations without forcing every kernel author to rewrite at the metal level — and that can decouple rapidly changing AI frontends from rapidly changing silicon.
Future Directions
- Extending beyond the two case studies. The paper demonstrates matrix multiplication and one attention block; whether the approach scales to full transformer stacks or a broader set of AI workloads is an open question in the presented material.
- Improving efficiency toward hand-optimized quality. The reported figure of up to 78.7% compute efficiency leaves room to close the remaining gap with hand-tuned implementations, and to understand which workloads close it easily.
- Generalizing to more spatial accelerators. The dialect is explicitly designed to be vendor-agnostic, and the stated intent is to target a wide range of spatial accelerators beyond AMD's NPU — a claim that needs demonstration across other hardware.
- Scheduling policy and the compiler/runtime split. The design deliberately divides scheduling responsibility between compiler and runtime, with the compiler able to hand off multiple variants of a launch (for example, different opportunistic granularities such as one column or the whole array). Understanding how to choose among those variants, and how much authority each side should hold, remains an area for further work.
- Herd sizing and backend feasibility. Because
air.herdoperations are scheduled atomically, users must be aware that architectures may cap the resources that can be guaranteed to run concurrently, and lowering can fail at backend compilation if an unimplementable herd is created — a practical usability constraint the design leaves open.
Target Audience
Compiler engineers and researchers working on MLIR, accelerator code generation, or domain-specific compilers will get the most from this paper, since the core contribution is an IR design and its lowering flow. Hardware architects interested in how spatial accelerators such as NPUs expose placement, memory hierarchy, and DMA control will find the hardware-trends analysis and the AIR abstraction mapping directly relevant. Performance engineers targeting AMD NPUs — or evaluating whether a higher-level flow can approach hand-tuned MLIR-AIE code — are the practical audience for the two case studies. Readers without a compiler background can still follow the motivation sections and the high-level results, but the dialect details require Intermediate-to-Advanced familiarity with compiler infrastructure.
Authors’ abstract
General-purpose compilers abstract away parallelism, locality, and synchronization, limiting their effectiveness on modern spatial architectures. As modern computing architectures increasingly rely on fine-grained control over data movement, execution order, and compute placement for performance, compiler infrastructure must provide explicit mechanisms for orchestrating compute and data to fully exploit such architectures. We introduce MLIR-AIR, a novel, open-source compiler stack built on MLIR that bridges the semantic gap between high-level workloads and fine-grained spatial architectures such as AMD's NPUs. MLIR-AIR defines the AIR dialect, which provides structured representations for asynchronous and hierarchical operations across compute and memory resources. AIR primitives allow the compiler to orchestrate spatial scheduling, distribute computation across hardware regions, and overlap communication with computation without relying on ad hoc runtime coordination or manual scheduling. We demonstrate MLIR-AIR's capabilities through two case studies: matrix multiplication and the multi-head attention block from the LLaMA 2 model. For matrix multiplication, MLIR-AIR achieves up to 78.7% compute efficiency and generates implementations with performance almost identical to state-of-the-art, hand-optimized matrix multiplication written using the lower-level, close-to-metal MLIR-AIE framework. For multi-head attention, we demonstrate that the AIR interface supports fused implementations using approximately 150 lines of code, enabling tractable expression of complex workloads with efficient mapping to spatial hardware. MLIR-AIR transforms high-level structured control flow into spatial programs that efficiently utilize the compute fabric and memory hierarchy of an NPU, leveraging asynchronous execution, tiling, and communication overlap through compiler-managed scheduling.