Research
Optimizing PyTorch Inference with LLM-Based Multi-Agent Systems
Overview Research area: LLM-based multi-agent systems applied to AI/ML systems optimization, specifically generating and tuning GPU kernels for PyTorch inference. Technical level: Advanced. The paper

- arXiv
- 2511.16964
- Published
- 2025-11-21
- Authors
- Kirill Nagaitsev, Luka Grbcic, Samuel Williams, Costin Iancu
AI summary
Overview
- Research area: LLM-based multi-agent systems applied to AI/ML systems optimization, specifically generating and tuning GPU kernels for PyTorch inference.
- Technical level: Advanced. The paper assumes familiarity with LLM agents, evolutionary search (mutation/crossover, islands, explore-exploit balance), GPU kernel programming in CUDA and Triton, and inference compilers such as TorchInductor and TensorRT.
- Scope: The paper builds a logical framework for comparing multi-agent PyTorch optimization systems and empirically ablates two implementations of that framework (PIKE-B and PIKE-O) on a refined KernelBench suite using H100 GPUs.
What This Paper Is About
Getting maximum performance out of AI inference on a GPU traditionally requires either hand-written custom kernels or model compilers, both of which are costly and slow to adapt as GPU hardware changes. Recent work shows LLM-based multi-agent systems can perform this optimization automatically, but the internal dynamics of such systems—how agents, prompts, and solution libraries should be configured—have not been systematically studied. This paper lays out a common logical framework for these systems and empirically determines which configurations (exploration-heavy versus exploitation-heavy, with or without error fixing) produce the best speedups on PyTorch models.
Key Contributions
- A logical framework for comparing multi-agent PyTorch optimization systems, decomposing the process into library, seed selection, prompt construction, evaluation, and post-processing stages, with parameter choices that push the system toward exploration or exploitation.
- Two implementations of that framework: PIKE-B, a multi-agent evolutionary branching strategy, and PIKE-O, an OpenEvolve-based strategy. Together with the framework they are collectively called PyTorch Inference Kernel Evolution (PIKE).
- An empirical finding that exploit-heavy strategies with error fixing substantially outperform explore-heavy strategies, and that performance correlates with the granularity of optimization steps—aggressive, exploit-heavy steps do better within the paper's budget.
- State-of-the-art speedups on a refined KernelBench suite (METR-refined, with the authors' own Level 3-pike filtering), with best solutions implemented in both CUDA and Triton. Code is publicly available at https://github.com/pike-project/pike.
Main Findings
- Best overall result: PIKE-B with the error-fixing agent (EFA) achieved the highest speedup on Level 3-pike at 2.88× over PyTorch Eager. The abstract reports an average 2.88× speedup over PyTorch Eager and 1.85× over torch.compile on an H100 across diverse KernelBench tasks.
- Error fixing is critical for PIKE-B: Removing the EFA dropped PIKE-B from 2.88 to 1.98 on Level 3-pike. The paper attributes this to the EFA identifying and correcting code generated by the Code Optimization Agent (COA).
- Brainstorming matters less than error fixing: PIKE-B with no Initial Brainstorming Agent (IBA) achieved 2.68, close to the full PIKE-B's 2.88.
- Cheap error fixing trades cost for reliability: PIKE-B with a cheap EFA (Gemini 2.5 Flash) reached 2.59 on Level 3-pike. It lowered mean EFA query cost across Level 3-pike tasks from $0.15 to $0.04, but the percentage of solutions working before the 5-attempt EFA maximum fell from 79.3% to 71.1%. At a $25-per-task budget, that cheap-EFA configuration gave the largest speedup, 2.51.
- Model choice matters: PIKE-B with OpenAI gpt-oss-120b reached 2.31 on Level 3-pike, and PIKE-B with o3-mini-high reached 1.58. Because METR's best solutions across tasks scored 1.40 using the same o3-mini-high model, the same 300-query budget, and the same H100 hardware, the authors argue PIKE-B is strictly better in that comparison (METR uses best-of-N across several models, each with a 300-query budget).
- Default PIKE-O is barely affected by error fixing: Default PIKE-O scored 2.17 with EFA and 2.15 without, suggesting its multi-island architecture and redundant seed diversity compensate for errors.
- Tuning PIKE-O toward exploitation helps a lot: The progression on Level 3-pike went 2.17 (default) → 2.10 (mutation-only) → 1.99 (mutation, no parallelism) → 2.75 (single island) → 2.81 (single island, exploit ratio 1, short-term library of 4 elites). PIKE-O (mut,npar,1isl,EO), which has the high exploit ratio but a larger library, reached only 2.38, suggesting the library-size reduction drives much of the gain.
- Level 5 results: PIKE-B reached 2.57 at full budget versus PIKE-O's 2.25; by cost, PIKE-B reached 2.44 versus PIKE-O's 2.2 at a $50-per-task budget. PIKE-O (mut,npar,1isl,EO,SL) was the best PIKE-O variant at 2.47 (full budget) and 2.33 ($50 budget). PIKE-O (mut,npar,1isl,EO) reached 2.41 at full budget and beat the former variant at the $50 budget. METR scored 1.50 on Level 5, the best of the non-PIKE competitors there.
- Competitors underperform PIKE under the same budgets: On Level 3-pike, torch.compile scored 1.64, TensorRT 1.41, and METR 1.40. On Level 5, they scored 1.29, 1.25, and 1.50 respectively.
- PIKE-B changes code more aggressively than PIKE-O: PIKE-O required fewer error fixes (most tasks averaging 1.5 attempts or less), while PIKE-B generally exceeded 2 attempts, and PIKE-B modified more lines of code per optimization step. PIKE-O makes smaller, more conservative changes that are easier to correct.
- Exploit-heavy optimization drives the need for error fixing: The authors conjecture that error-fixing is needed because exploit-heavy optimization rapidly increases code complexity, and that the exploitative PIKE-O variant behaves similarly to PIKE-B in error-fix effort and magnitude of code changes.
- Parallelism has an unintended exploratory bias: The authors note that high parallelism (parallel evaluations = 10) lets quick-to-evaluate but low-quality solutions push the algorithm toward early-stage seeds rather than maturing complex, higher-performing solutions.
Methodology in Plain English
The researchers first defined a general "logical framework" describing what any LLM-based PyTorch optimization system does: hold candidate solutions in a library, pick a seed solution, build a prompt, query an LLM to generate new code, evaluate it (compile, check correctness, measure speed), optionally loop through an error-fixing agent, then store the result. Parameters determine whether the system mostly explores new solutions or exploits the best ones found so far.
They then built and compared two implementations. PIKE-B is an evolutionary branching search: it generates 10 initial ideas, evaluates them in parallel, fixes errors, takes the top 4 solutions, and duplicates them across a population of 10 for the next round—100% exploit-based, mutation-only, short-term memory, no islands. PIKE-O builds on the open-source OpenEvolve framework, which by default uses islands, crossover from five or more elites, long-term memory, and parallel evaluations; the authors added error fixing through the EFA, which OpenEvolve originally lacked, and ran a sequence of ablations stepping it toward PIKE-B's exploit-heavy behavior.
For evaluation they used a refined KernelBench suite: Level 3-pike (30 curated component tasks averaging 85 lines of code) and METR's Level 5 (14 frontier tasks averaging 493 lines of code, the largest being HunyuanTransformer at 1,268 lines). Solutions were compiled and checked with the same numerical equivalence tests and tolerances as prior work, run in Docker containers on bare-metal NVIDIA H100 GPUs with 80 GB HBM3 memory over PCIe, with Intel Xeon Platinum 8480+ CPUs and 40+ CPU threads allowing 20+ parallel compilations. Timing used Triton's do_bench with at least one warmup run and the mean of many timed runs, under a GPU lock. Each run got a 300-query budget across all agents, matching METR's budget, using Gemini 2.5 Pro, with ablations using Gemini 2.5 Flash, OpenAI gpt-oss-120b, and o3-mini-high. Speedups were computed relative to PyTorch Eager, clamped to 1 if below 1, and reported as geometric means per level.
Why This Matters
- Impact on research: The paper shifts attention from "does an LLM agent system work?" to "which configuration of a multi-agent system works, and why?"—providing an ablation methodology and open source code that other groups can reuse to compare agent strategies rather than only reporting end-to-end speedup numbers.
- Reducing reliance on hand-tuned kernels: Fully automated optimization can lower the barrier for teams that cannot afford GPU kernel specialists, who must otherwise master GPU parallelism, memory hierarchies, and scheduling.
- Practical deployment decisions: The cost-vs-speedup curves let teams pick a configuration based on budget—for example, exploit-heavy PIKE-B with a cheap EFA at roughly $25 per task, or the more capable model when accuracy of fixes matters more than cost.
- Frontier model workloads: Level 5 tasks drawn from DeepSeek-V3, Llama 3, RWKV, SD3, Mamba-2, S4, and Hunyuan Video represent architectures where new kernel generation is most valuable.
- Industry relevance: Because the systems are built on PyTorch, CUDA, and Triton inputs and compared against torch.compile and TensorRT, the results speak directly to production inference stacks. The value proposition is meaningful across workloads: the paper notes that the original custom FlashAttention kernel achieved a 7.6× speedup over a generic PyTorch implementation, illustrating what hand-written kernels can deliver, and that inference optimization must continually adapt as GPUs evolve.
Future Directions
- Complete hyperparameter sweeps: The authors state that the high cost of experimentation prevented a complete hyperparameter sweep for both methods within their budget, leaving open the question of how configurations interact across the full parameter space.
- Scaling budget and hardware: The paper sweeps per-task budget from zero up to a maximum, but the frontier of much larger budgets, additional GPUs, or newer hardware targets is not explored.
- Model evolution: Results are tied to Gemini 2.5 Pro/Flash, gpt-oss-120b, and o3-mini-high; how the exploit-heavy conclusion holds for future LLMs is an open question.
- Understanding the error-fixing dependency: The authors present evidence—and conjecture—that the need for error fixing is driven by exploit-heavy optimization rapidly increasing code complexity, but they did not run the PIKE-O exploitative variant without EFA due to budget constraints, so this hypothesis remains unconfirmed.
Target Audience
This paper is most valuable to AI systems researchers and engineers who build or evaluate LLM-agent pipelines for code generation and performance optimization, particularly those working on GPU kernel generation, inference serving, or compilers. It is also relevant to practitioners selecting between automated optimization approaches and hand-tuning, and to readers interested in how explore-exploit tradeoffs play out in concrete multi-agent search systems. A background in GPU programming, PyTorch, and evolutionary search is helpful for following the ablation reasoning in detail.
Authors’ abstract
Maximizing performance on available GPU hardware is an ongoing challenge for modern AI inference systems. Traditional approaches include writing custom GPU kernels and using specialized model compilers to tune high-level code for specific GPU targets. Recent work shows that LLM-based multi-agent systems can effectively perform such tuning, often outperforming existing compilers and eliminating the need for manual kernel development. However, the dynamics of multi-agent systems for this task remain unexplored. In this work, we present a logical framework for comparing multi-agent PyTorch optimization systems. Our evaluation shows that exploit-heavy strategies perform best when paired with error-fixing agents, and that performance correlates with the granularity of optimization steps. The best implementation achieves an average 2.88x speedup over PyTorch Eager (1.85x over torch.compile) on an H100 GPU across diverse tasks in KernelBench, a benchmark suite covering a range of machine learning architectures in PyTorch. Code is publicly available at: https://github.com/pike-project/pike