Skip to content
AI.info

Research

ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization

ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization Overview Research area: Machine learning for compilers — specifically LLM-guided compiler auto-tuning and the phase-ordering problem (c

ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
arXiv
2602.00087
Published
2026-01-23
Authors
Haolin Pan, Lianghong Huang, Jinyuan Dong, Mingjie Xing, Yanjun Wu

AI summary

ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization

Overview

Research area: Machine learning for compilers — specifically LLM-guided compiler auto-tuning and the phase-ordering problem (choosing and ordering LLVM optimization passes to minimize runtime cycles).

Technical level: Advanced. The paper assumes familiarity with LLVM intermediate representation (IR), optimization passes, genetic algorithms, supervised fine-tuning, and reinforcement learning (GRPO).

Scope in one sentence: The paper proposes ECCO, a framework that trains an LLM on reverse-engineered causal evidence to act as a high-level "strategist" whose optimization intents steer a genetic algorithm, reporting an average 24.44% cycle reduction over LLVM opt -O3 across seven benchmark suites.

What This Paper Is About

Compiler auto-tuning is hard because the space of pass orderings is combinatorial and non-convex. Traditional search methods (genetic algorithms, Bayesian optimization, reinforcement learning) treat the compiler as a black box and optimize a scalar metric without understanding why a pass helps, while existing LLM approaches tend to do superficial pattern matching on code-and-flag pairs without grasping the causal mechanism by which a pass changes code and improves performance. ECCO's goal is to close this gap by teaching a model the causal chain linking static code features to structural transformations to performance gains, and then pairing that semantic reasoning with a genetic algorithm that handles the fine-grained combinatorial search the LLM cannot do reliably on its own.

Key Contributions

  1. Evidence-driven causal paradigm. The authors propose a reverse-engineering pipeline that prunes optimal pass sequences down to critical passes and extracts verifiable evidence (IR structural diffs, Autophase feature deltas, per-pass cycle gains, and pass-order synergy signals), then distills this into a Chain-of-Thought training corpus via "simulated predictive reasoning."
  2. Collaborative Strategist–Tactician framework. A collaborative inference scheme in which the LLM defines high-level optimization intents that guide — but do not dictate — the mutation operator of a genetic algorithm, using a soft probabilistic bias rather than hard pruning of the search space.
  3. Two new datasets for the community. A forensic evidence dataset mapping optimization passes to verifiable IR feature deltas, and a chain-of-thought reasoning dataset for training interpretable optimization policies.
  4. Empirical evaluation on seven benchmark suites. ECCO is compared against traditional search heuristics and direct LLM prompting, and its rationale faithfulness is audited by five LLM judges.

Main Findings

  • Average cycle reduction over LLVM opt -O3: 24.44% across the seven benchmark suites (blas, cbench, chstone, mibench, npb, opencv, tensorflow), with per-suite results of 12.99%, 35.19%, 35.50%, 27.12%, 32.97%, 15.58%, and 11.72% respectively, using Qwen2.5-7B with Best-of-32 sampling. Largest gains are on cbench, chstone, and mibench.
  • Beats the strongest traditional baselines: ECCO's 24.44% average exceeds CFAST (23.70%), GRACE (23.14%), CompTuner (22.88%), PDCAT (22.79%), GA (22.14%), OpenTuner (22.14%), RIO (21.98%), and TPE (19.88%).
  • Direct LLM prompting underperforms even with Best-of-32: Qwen3-Coder reaches 16.25%, Kimi-K2 16.61%, DeepSeek-V3.2 16.06%, and GPT5-chat 16.56% on average — all failing to surpass TPE's 19.88%. The authors interpret this as a capability ceiling for general-purpose reasoning on phase ordering.
  • Pruning reveals heavy redundancy in search-derived sequences: Iterative greedy pruning reduced the average sequence length from 18.5 to 4.5 passes — a 73.7% reduction — and for over 60% of programs the redundancy exceeded 70%.
  • Both evidence and chain-of-thought matter, and structure alone beats blind imitation: With the genetic algorithm removed entirely, the ablated "w/o CoT" variant drops to 1.27% average improvement at N=1 greedy decoding (16.25% at N=32), and "w/o Evidence" reaches 3.52% (17.55% at N=32), versus standard ECCO-1.5B (w/o GA) at 4.38% (19.02% at N=32).
  • Larger models do not automatically win: ECCO-7B (w/o GA) peaks at 19.43% at N=32, only slightly above ECCO-3B (19.13%) and ECCO-1.5B (19.02%); the 7B model regresses at N=1 (5.04%), which the authors attribute to overfitting/memorization of training noise. They describe the 3B model as a "sweet spot."
  • A measurable "tactical gap" of roughly 5%: The standalone LLM policy saturates near 19.4% even with 7B parameters and extensive sampling, versus 24.44% for the full collaborative system. Replacing the GA with an iterative LLM refiner (Appendix E) saturated at about 19.36%, corroborating that the last mile of pass ordering requires combinatorial search.
  • High rationale faithfulness under LLM-as-judge auditing: Across five judges, average accuracy was 78.96% (DeepSeek-V3.2), 82.89% (Claude-4.5-Sonnet), 96.47% (Qwen3-Coder), 97.08% (GPT5-chat), and 99.84% (Kimi-K2). The table reports a consensus average of 91.05%, while the analysis text states the consensus averages 91.24%; cbench-v1 and chstone-v0 reach 100% under every judge.
  • Complexity-dependent faithfulness: Lower-scoring datasets under strict judges are blas-v0 (84.83% consensus), opencv-v0 (83.13%), and tensorflow-v0 (82.67%), which the authors attribute to complex vectorization and memory patterns that Autophase features describe less well.

Methodology in Plain English

The pipeline has three stages.

Stage 1 — Building a causal dataset. The authors start from high-performance pass sequences found with the CFSAT framework, which tend to be over-provisioned. They apply an iterative greedy pruning algorithm: try removing each pass in turn, keep the removal if cycles do not get worse, and repeat until nothing more can be removed. This isolates the passes that actually matter. For each remaining step, they run a forensic reconstruction, capturing three kinds of evidence: the textual IR diff before and after the pass (structural evidence), the shift in the 56-dimensional Autophase feature vector (feature evidence), and the marginal cycle improvement of that step (performance evidence). They also swap adjacent passes to detect ordering constraints ("positive synergy"). Because dynamic measurements are unavailable for unseen programs at inference time, they prompt a teacher model (Claude-4.5-Sonnet) to convert the evidence into rationales — with a "feigned reasoning" constraint that forces the narrative to look like it was derived purely from the initial static features, presenting the ground-truth deltas as the model's own predictions. The target model is then fine-tuned on these (initial features, rationale + sequence) pairs so it learns to simulate intermediate states internally.

Stage 2 — Two-stage policy optimization. The policy is initialized with Qwen2.5-Instruct and first aligned by supervised fine-tuning to establish a reasoning prior and a parseable output format. It is then optimized with GRPO reinforcement learning using two rewards: a discrete format reward requiring distinct <think> and <answer> tags with a syntactically valid JSON array of LLVM pass flags, and a continuous performance reward equal to a scaled relative speedup against the LLVM -O3 baseline.

Stage 3 — Collaborative inference. Before deployment, the authors run a large-scale ablation over the training corpus to measure the marginal benefit of removing each pass from its optimal sequence, aggregate this into a Global Expected Benefit per pass, and designate top-k "star passes" per functional category (Loop, Scalar, Vectorization). At inference, the LLM predicts an initial sequence and an intent distribution over optimization categories. A genetic algorithm is seeded with that sequence and mutates it using a probability that mixes the LLM's intent with uniform exploration — a soft bias rather than a hard restriction of the search space — so any pass retains non-zero probability and the search can theoretically recover from strategic mistakes.

Evaluation. The training corpus is built from the CompilerGym corpus following GRACE, yielding 9,327 samples. Evaluation uses seven benchmark suites. Optimization effectiveness is measured as the relative cycle reduction over -O3, with cycles estimated by llvm-mca. Models are Qwen2.5-Instruct at 1.5B, 3B, and 7B; experiments run on Intel Xeon Gold 6430 servers with 4× NVIDIA H100 GPUs, using LLVM version 10.0.0.

Why This Matters

Impact on research. The paper argues that the performance gap between general-purpose LLMs and domain-aligned models on compiler tasks is not closed by scaling general reasoning — GPT-5 performing similarly to Qwen3 is presented as evidence that explicit alignment with IR semantics is a prerequisite. It also frames interpretability as a measurable capability rather than a byproduct, offering an auditable evaluation protocol for optimization rationales. The release of two datasets — a forensic evidence dataset and a Chain-of-Thought reasoning dataset — supports reproducibility and follow-on work.

Real-world applications:

  • Compiler and runtime teams tuning optimization pipelines for specific hardware targets, where a fixed -O3 sequence does not exploit workload characteristics.
  • Performance engineering for latency-sensitive code such as numerical kernels, embedded workloads, and media processing, where cycles directly determine cost or feasibility.
  • Build-system and toolchain vendors that need optimized binaries without hand-tuning per program, using the trained policy as an automated pass-ordering advisor.
  • Audit and compliance contexts — such as safety-critical or regulated compilation pipelines — where the ability to produce a human-readable, evidence-checkable justification for a transformation sequence matters as much as the speedup.

Industry relevance. The work targets a practical bottleneck in toolchain optimization and reports gains relative to a widely used LLVM baseline, which makes the results directly comparable to existing deployment practices. The hybrid architecture is also pragmatic from an engineering standpoint: it uses small models (1.5B–7B) rather than frontier-scale ones, and the LLM-only mode at Best-of-32 approaches traditional baselines, meaning partial deployment without the genetic algorithm is still viable. The paper does not report tuning-time or compilation-overhead costs, so the wall-clock tradeoff against -O3 is not quantified here.

Future Directions

  • Closing the tactical gap. The roughly 5% gap between the standalone LLM (~19.4%) and the collaborative system (24.44%) is presented as evidence that a hybrid is necessary; whether a different refinement mechanism or a search method other than a GA can be more efficient at the last mile remains open.
  • Improving rationale faithfulness on complex workloads. opencv and tensorflow score lowest under strict judges (83.13% and 82.67% consensus), and the authors attribute this to Autophase features being less descriptive for vectorization and memory patterns — suggesting richer or alternative feature representations are needed.
  • Removing the LLM-as-judge circularity. The authors explicitly acknowledge that using LLMs to evaluate LLMs risks bias and leniency; they mitigate this with a five-judge ensemble, but an evidence-grounded, non-LLM faithfulness metric would be a stronger foundation.
  • Understanding the capacity–generalization tradeoff. The 7B model underperformed the 3B model at N=1 and converged with smaller models at N=32, which the authors attribute to overfitting; the relationship between model scale, training budget, and generalization in this domain is not resolved.

Target Audience

This paper is best suited to researchers and practitioners in compiler engineering, systems optimization, and machine learning for code: people working on LLVM pass pipelines, auto-tuning frameworks, or applying LLMs to structured program-analysis tasks. It is also relevant to readers interested in reinforcement learning with verifiable rewards, evidence-grounded chain-of-thought distillation, and hybrid neuro-symbolic search. It requires comfort with compiler IR concepts and RL fine-tuning terminology; readers without that background will find the introduction and discussion accessible but the methodology dense.

Authors’ abstract

Compiler auto-tuning faces a dichotomy between traditional black-box search methods, which lack semantic guidance, and recent Large Language Model (LLM) approaches, which often suffer from superficial pattern matching and causal opacity. In this paper, we introduce ECCO, a framework that bridges interpretable reasoning with combinatorial search. We first propose a reverse engineering methodology to construct a Chain-of-Thought dataset, explicitly mapping static code features to verifiable performance evidence. This enables the model to learn the causal logic governing optimization decisions rather than merely imitating sequences. Leveraging this interpretable prior, we design a collaborative inference mechanism where the LLM functions as a strategist, defining optimization intents that dynamically guide the mutation operations of a genetic algorithm. Experimental results on seven datasets demonstrate that ECCO significantly outperforms the LLVM opt -O3 baseline, achieving an average 24.44% reduction in cycles.

Read the original paper