Research
Pie: A Programmable Serving System for Emerging LLM Applications
Overview Research area: Systems for machine learning — specifically LLM inference serving infrastructure, sitting at the intersection of operating systems, programming-language/runtime design, and nat

- arXiv
- 2510.24051
- Published
- 2025-10-28
- Authors
- In Gim, Zhiyao Ma, Seung-seob Lee, Lin Zhong
AI summary
Overview
Research area: Systems for machine learning — specifically LLM inference serving infrastructure, sitting at the intersection of operating systems, programming-language/runtime design, and natural language processing.
Technical level: Advanced. The paper assumes familiarity with KV caches, prefill-decode loops, continuous batching, and GPU kernel execution.
Scope: The paper presents Pie, a programmable LLM serving system that replaces the conventional monolithic generation loop with granular API handlers controlled by user programs called inferlets, and evaluates whether this design matches standard-task performance while improving latency and throughput on emerging agentic and reasoning workloads.
What This Paper Is About
Existing LLM serving systems such as vLLM and TGI are built around a single, fixed prefill-decode loop with system-wide policies for KV cache management, one closed token prediction-and-sampling procedure, and no native way to interleave inference with external computation or I/O. Pie's goal is to hand end-to-end control of generation to application-supplied programs, so that an application can define its own KV cache strategy, its own decoding logic, and its own mix of inference, computation, and network calls — without modifying the serving system itself.
Key Contributions
-
A diagnosis of three limitations in current serving infrastructure: implicit, system-wide KV cache management (motivating requirement R1); an inflexible predict-then-sample decoding loop with no per-request customization (R2); and poor integration of token generation with external tools, API calls, and code execution (R3).
-
The Pie architecture, which dismantles the monolithic generation pipeline into fine-grained, independent service handlers (for example embedding, KV cache allocation/update, forward pass, and sampling) exposed through an API, organized into a three-layer system: application, control, and inference.
-
A Turing-complete programming model for inferlets, providing 42 APIs — 18 dedicated to LLM execution in the inference layer and 24 for runtime management, inter-inferlet communication, and I/O — that give applications full, end-to-end control over the generative workflow, KV cache, and I/O. LLM-related APIs are organized into traits (Allocate, Forward, InputText, InputImage, Tokenize, OutputText) to keep the system model-agnostic and extensible.
-
An implementation and evaluation in which a diverse set of LLM techniques are implemented as inferlets (attention variants, constrained and speculative decoding, deliberate prompting strategies, and agentic workflows), showing performance matching state of the art on standard tasks and improving latency and throughput on emerging applications. Pie is open-sourced at https://github.com/pie-project/pie.
Main Findings
-
Standard-task overhead is modest: Pie matches state-of-the-art performance on standard tasks, with 3–12% latency overhead for text completion.
-
Agentic and reasoning workflows improve substantially: Pie reports 1.1×–2.4× lower latency and 1.3×–3.4× higher throughput on Graph-of-Thought and agentic workflows by enabling application-specific optimizations. (The abstract states the throughput range as 1.3×–3.4× higher.)
-
Agentic latency and throughput figures: Using a 1B model with 8 (ReACT), 8 (CodeACT), and 32 (Swarm) external I/Os per agent, measured latencies are 4.27 s, 3.18 s, and 6.14 s, with throughputs of 29.94, 40.18, and 5.21 agents/s respectively. Pie reduces latency by up to 15% and increases throughput by up to 30% on ReACT.
-
I/O ratio drives the gain: The paper states that the performance gains are closely tied to the ratio of I/O to the total number of tokens in the workflow. The provided content is truncated mid-sentence at this point, so the remaining analysis of that relationship is not reported here.
-
Applications are cheap to express: Implemented inferlets range from 38 lines of code (text completion, 129 KB Wasm binary) to 255 lines (speculative decoding, 152 KB) and 225 lines (EBNF decoding, 2 MB). Agent applications range from 60 lines (Agent-ReACT, 309 KB) and 62 lines (Agent-CodeACT, 6.7 MB, larger due to an embedded JavaScript runtime) to 95 lines (Agent-SWARM, 135 KB). The support library itself is 728 lines.
-
A native inference layer is faster: A native C++/CUDA implementation of the inference layer that avoids PyTorch achieves 10–30% lower end-to-end latency and more efficient GPU memory utilization than the Python/PyTorch counterpart, owing to custom memory management that pre-allocates and reuses GPU buffers. Because it currently supports only a subset of API traits (Forward, InputText), the authors omit a full evaluation of it.
-
Context for the KV cache argument: The paper notes that implementing fine-grained KV cache features in a monolithic architecture can require invasive changes, citing that support for beam search was considered for removal from vLLM (v0.6.3) due to its complexity.
Methodology in Plain English
The authors began by identifying what modern LLM applications need that existing servers cannot give them, and grouped those needs into three requirements: application-specific KV cache control, customizable generation processes, and integrated computation and I/O. They then rebuilt the serving stack around a different division of labor.
Instead of one hard-coded generation loop, the system exposes small, well-defined operations — embed text or images into embeddings, allocate and manipulate paged KV cache blocks, run a forward pass, and obtain a next-token distribution — as API calls. An application writes an inferlet, a program that calls these operations in whatever order and combination it wants, effectively reimplementing the generation loop itself. Inferlets are compiled to WebAssembly and executed in a sandboxed runtime (the authors use wasmtime, with WASI providing interfaces such as network I/O), so applications written in any language that targets Wasm, such as C++, Rust, or Python, can be deployed. The system is split into an application layer (managing inferlet lifecycles and isolation), a control layer (handling non-GPU calls, virtualizing resources, and batching GPU-bound calls), and an inference layer (executing batched GPU work). Because hundreds of inferlets can run concurrently with different strategies, the control layer's batch scheduler groups compatible calls both vertically (consecutive same-type calls within one command queue) and horizontally (same-type calls across queues), using a work-conserving policy to decide when to dispatch.
For evaluation, the authors implemented a broad table of LLM techniques as inferlets — tree-of-thought, recursion-of-thought, graph-of-thought, skeleton-of-thought, prefix caching, modular caching, EBNF decoding, beam search, watermarking, output validation, speculative decoding, Jacobi decoding, attention sink, windowed attention, hierarchical attention, and three agents (ReACT, CodeACT, SWARM) — and compared against vLLM (v0.6.0), SGLang (v0.4.4), LMQL (v0.7.3), and StreamingLLM where relevant, replicating the same high-level logic in Python scripts against each baseline's API server. All three of Pie, vLLM, and SGLang use the FlashInfer GPU backend to isolate architectural differences from kernel-level optimization differences. Measurements were taken on a GCP G2 instance (g2-standard-32) with an NVIDIA L4 GPU with 24 GB of memory, using Llama 3 models (1B, 3B, 8B) in BF16, with a remote Python client on a campus network.
Why This Matters
Impact on research. The paper reframes LLM serving as a programmability problem rather than purely a batching-and-kernel-optimization problem. By arguing that all existing systems effectively run a single, fixed inferlet (an autoregressive loop), it gives systems researchers a vocabulary for comparing serving designs and a substrate for experimenting with new decoding and caching strategies without patching a memory manager or scheduler.
Real-world applications:
- Agentic assistants that interleave generation with web API calls, as in the ReACT inferlet, avoiding client round-trips and re-prefill of interaction history.
- Code-writing and tool-using agents that execute code inside the generation flow, as in the CodeACT inferlet, which embeds a JavaScript runtime.
- Multi-agent systems where agents communicate with one another, as in the SWARM inferlet, which uses message-passing APIs rather than external orchestration.
- Structured and constrained output for applications needing schema-conformant or grammar-constrained text, as in the EBNF decoding inferlet, evaluated against LMQL.
- Long-context and streaming use cases, where attention sink and windowed attention inferlets manage the KV cache explicitly.
Industry relevance. The design targets the parts of an LLM stack that provider teams currently fork or patch — cache eviction, decoding, and tool integration. Explicit resource allocation and deallocation through APIs, per-inferlet virtual address spaces, import/export of KV pages between inferlets, and priority-hinted command queues suggest a path toward multi-tenant serving where different applications run different strategies on shared hardware. The overhead figures reported (3–12% on standard text completion) are the kind of number a production team would weigh against the flexibility gained, and the open-source release plus the documented artifact badges (artifacts available, functional, results reproduced) support reproducibility.
Future Directions
-
Extending the trait system to new modalities and capabilities. The paper suggests that adding audio input or other capabilities requires only defining a new trait and having models implement it, and proposes an
IntrospectiveForwardtrait returning internal statistics such as token-level attention scores, which would let inferlets implement schemes like PyramidKV and AdaKV. -
Completing and evaluating the native inference layer. The C++/CUDA implementation currently supports only a subset of API traits (Forward, InputText) and was therefore not fully evaluated; broadening trait coverage and measuring it end-to-end is a clear next step.
-
Reducing the programming complexity the model introduces. The API deliberately prioritizes fine-grained programmability over simplicity, and the authors mitigate this with a Rust support library, procedural macros, and higher-level abstractions such as
Context. How far those abstractions can go before flexibility is lost is left open. -
Model coverage beyond the Llama family. The API handlers currently support Llama-family models only, so generalization to other architectures — and the effect of that generalization on the batching and resource-management layers — remains an open question.
Target Audience
This paper is most useful to systems researchers and engineers who build or operate LLM serving infrastructure — particularly those working on KV cache management, decoding-strategy integration, or agent/tool orchestration layers. It is also relevant to researchers developing novel reasoning and decoding techniques who currently need invasive changes to a serving system to deploy them, and to graduate students in systems, operating systems, or MLSys who want a concrete example of applying sandboxing, virtualization, and batching abstractions to LLM inference. Readers without background in inference serving will need to be comfortable with prefill-decode mechanics, paged attention, and GPU-batching terminology.
Authors’ abstract
Emerging large language model (LLM) applications involve diverse reasoning strategies and agentic workflows, straining the capabilities of existing serving systems built on a monolithic token generation loop. This paper introduces Pie, a programmable LLM serving system designed for flexibility and efficiency. Pie decomposes the traditional generation loop into fine-grained service handlers exposed via an API and delegates control of the generation process to user-provided programs, called inferlets. This enables applications to implement new KV cache strategies, bespoke generation logic, and seamlessly integrate computation and I/O-entirely within the application, without requiring modifications to the serving system. Pie executes inferlets using WebAssembly, benefiting from its lightweight sandboxing. Our evaluation shows Pie matches state-of-the-art performance on standard tasks (3-12% latency overhead) while significantly improving latency and throughput (1.3x-3.4x higher) on agentic workflows by enabling application-specific optimizations.