Skip to content
AI.info

Research

LLM-Driven Composite Neural Architecture Search for Multi-Source RL State Encoding

Overview Research area: Reinforcement learning (RL) state representation learning, neural architecture search (NAS), and LLM-guided automated machine learning. Technical level: Advanced. The paper ass

LLM-Driven Composite Neural Architecture Search for Multi-Source RL State Encoding
arXiv
2512.06982
Published
2025-12-07
Authors
Yu Yu, Qian Xie, Nairen Cao, Li Jin

AI summary

Overview

Research area: Reinforcement learning (RL) state representation learning, neural architecture search (NAS), and LLM-guided automated machine learning.

Technical level: Advanced. The paper assumes familiarity with RL training loops (PPO), deep encoder architectures (CNNs, Transformers, GRUs, FFNs), and NAS terminology, and it builds on a formal optimization formulation over composite module design choices.

Scope: The paper introduces LACER, an LLM-driven pipeline that searches over composite state-encoder architectures (one module per input source plus a fusion module) for multi-source RL, and evaluates it on an RL-based mixed-autonomy traffic control benchmark.

What This Paper Is About

Reinforcement learning agents in real domains often receive observations from several heterogeneous sources at once — sensor measurements, time-series signals, images, and text instructions — each of which needs its own encoder, plus a module that fuses their outputs. Existing NAS methods mostly target single-modality supervised learning and ignore useful signals about how good each intermediate module's representation is. The paper formalizes composite NAS for RL state encoding and proposes an LLM-driven search pipeline that uses both task performance and representation-quality signals to find better encoders with fewer evaluations.

Key Contributions

  1. The paper introduces and formally defines the problem of composite NAS for state encoding in RL with multiple information sources, where source-specific modules and a fusion module are jointly optimized over the Cartesian product of their design choices.
  2. It proposes LACER, an LLM-driven NAS pipeline that uses language-model priors to guide the search using side information about module representation quality, rather than task performance alone.
  3. It describes how side information is computed and fed back to the LLM: mutual information (I(X;Y) = H(X) − H(X|Y)) and redundancy (R(X;Y) = H(X) + H(Y) − H(X,Y)) across feature pairs (before/after each source encoder, and between encoded source features and the fused features), plus average reward and the task metric.
  4. It instantiates and evaluates LACER on a mixed-autonomy traffic control task with 8 random seeds, reporting improved search efficiency and RL performance versus expert-designed encoders, DARTS, ENAS, PEPNAS, and the LLM-based GENIUS framework. The paper also describes search spaces for two further benchmarks (MiniGrid goal-oriented tasks and ManiSkill robotic control) in appendices, presented as examples rather than as reported results.

Main Findings

  • Both LACER variants beat all baselines: LACER-1 (one candidate per iteration) and LACER-5 (five candidates per iteration) "significantly outperform" the expert-designed architecture, the traditional NAS baselines (DARTS, ENAS, PEPNAS), and the LLM-based GENIUS baseline, measured as average traffic speed of the best architecture evaluated so far against number of evaluated candidates. The paper does not report the numeric average traffic speed values in the text provided.
  • Richer feedback signals are essential: Ablations removing the feature information (FI), the average reward (RI), and the initial architecture evaluation (IE) show that the original LACER-1 achieves the best performance; removing any component causes significant deterioration.
  • LLM query cost is negligible: Query time accounts for 1% of LACER's overall time cost, while evaluation time accounts for over 97% of the overall time cost of all methods. This analysis was run on an Intel Core i7 CPU and an NVIDIA RTX 4070 GPU to avoid node scheduling overhead.
  • Model and temperature choice matters: Testing Claude Sonnet 4.0 and GPT-4 at temperatures 0.0 and 1.0, LACER achieved optimal performance with Claude Sonnet 4.0 at temperature 1.0, which was adopted in the main experiments.
  • Search space sizes: The mixed-autonomy traffic control composite encoder spans approximately 26 million possible architectures; the MiniGrid composite encoder spans approximately 19 million. A search-space size for the ManiSkill encoder is not reported.
  • Training budget calibration: On the traffic benchmark, baseline RL analysis over 1M steps showed average reward converging around 200k steps and average speed exhibiting periodicity of roughly 25k steps, motivating 200k training steps and 50k evaluation steps per candidate. For MiniGrid, reshaped return converged around 1M steps and average return around 100 steps, motivating 1M training steps and 100 evaluation steps.

Methodology in Plain English

The researchers treat encoder design as a search problem and put a large language model in the role of a "neural architecture design agent." The pipeline starts from an expert-designed architecture. Each round, the LLM is given a pruned summary of the conversation so far — the task description, the search space definition, previous architectures, and their measured performance — and it proposes one or several new composite architectures. Responses are parsed automatically: the LLM is asked to mark its architecture description with a fixed prefix (for example, "New Architecture"), and regular expressions map the tokenized text into concrete design choices such as attention heads, dimension, expansion ratio, and depth.

Each proposed architecture is then trained end-to-end with the RL agent (PPO) for a fixed number of interaction steps and evaluated. Crucially, the feedback sent to the LLM at the next iteration is not just the task metric (average traffic speed) but also the average reward, which reflects RL convergence behavior, and feature information — mutual information and redundancy measured at the boundaries of each module and at the fusion output — which acts as a direct proxy for representation quality. The policy network architecture stays fixed; only the encoder modules change. The loop repeats until the evaluation budget is used up.

For the traffic benchmark, the encoder has four modules: a traffic encoder (FFN over a fixed-dimensional vector), a time encoder and a sequence encoder (Transformer-based, with multi-head self-attention for the two time-series inputs), and a fusion encoder (FFN over concatenated outputs). For fairness, every method is evaluated on 50 candidates total — 10 iterations for batch methods that propose five candidates per iteration, and 50 iterations for methods that propose one. Each experiment is repeated with 8 random seeds, and results are reported as means with error bars of two times the standard error across seeds. Experiments ran on the NYU Greene high-performance computing cluster with one NVIDIA GPU per run (Quadro RTX 8000 with 48 GB or Tesla V100 with 32 GB), 16 CPU cores, and 32 GB of RAM.

Why This Matters

Impact on research. The paper reframes encoder design for multi-source RL as a composite architecture search problem, in which submodules are themselves design variables rather than fixed functions — a setting the authors contrast with prior work on function networks with partial evaluations (Buathong et al.). It also positions representation-quality signals (mutual information, redundancy) as first-class search feedback, distinguishing it from LLM-based NAS methods such as GENIUS, LLMatic, LAPT-NAS, and SEKI, which the paper says are primarily designed for single-modality supervised tasks and use only a performance metric to update beliefs.

Real-world applications:

  • Mixed-autonomy traffic control, where connected autonomous vehicles share roads with human-driven vehicles and must combine temporal traffic evolution, current lane-level traffic state, and vehicle sequence history.
  • Goal-oriented navigation and instruction-following agents, which combine image observations with text instructions (the MiniGrid setting described in the appendix).
  • Robotic manipulation, which combines RGB perception with proprioceptive or other information signals (the ManiSkill setting described in the appendix).
  • General multi-sensor control systems, where sensor streams, time-series telemetry, and visual inputs must be encoded and fused before policy learning.

Industry relevance. Because RL architecture evaluation dominates cost — evaluations account for over 97% of total runtime across all methods while LLM query time is only 1% — any method that finds better encoders within a fixed evaluation budget translates directly into compute savings for teams training RL agents on multi-source observation data.

Future Directions

  • Applying LACER to broader applications such as goal-oriented tasks and robotic control with visual, textual, and sensor inputs, as stated in the conclusion.
  • Determining whether the mutual-information and redundancy signals transfer as effectively to image and text encoders as they do to the Transformer- and FFN-based traffic encoders.
  • Establishing how sensitive the pipeline is to the underlying LLM and its temperature settings beyond the two models (Claude Sonnet 4.0, GPT-4) and two temperatures (0.0, 1.0) tested.
  • Quantifying how search cost and benefit scale with the number of input sources, given that the composite search space represents a Cartesian product over all module design choices.

Target Audience

This paper is most useful for reinforcement learning researchers working on representation learning and sample-efficient training, NAS researchers interested in LLM-guided search and composite or multi-modal architecture spaces, and applied machine learning engineers building agents that consume heterogeneous observation sources. Readers need a working background in RL training procedures and neural architecture design to follow the methodology; the appendix material on search spaces and baseline alignment (mapping accuracy to average speed or average return, and sample size to training steps) will be particularly relevant to practitioners trying to reproduce or adapt the setup.

Authors’ abstract

Designing state encoders for reinforcement learning (RL) with multiple information sources -- such as sensor measurements, time-series signals, image observations, and textual instructions -- remains underexplored and often requires manual design. We formalize this challenge as a problem of composite neural architecture search (NAS), where multiple source-specific modules and a fusion module are jointly optimized. Existing NAS methods overlook useful side information from the intermediate outputs of these modules -- such as their representation quality -- limiting sample efficiency in multi-source RL settings. To address this, we propose an LLM-driven NAS pipeline in which the LLM serves as a neural architecture design agent, leveraging language-model priors and intermediate-output signals to guide sample-efficient search for high-performing composite state encoders. On a mixed-autonomy traffic control task, our approach discovers higher-performing architectures with fewer candidate evaluations than traditional NAS baselines and the LLM-based GENIUS framework.

Read the original paper