Research
InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-Manipulation
Overview Research area: Robot learning for humanoid loco-manipulation, combining behavioral foundation models (forward-backward representations), large language model driven reward design, and test-ti

- arXiv
- 2610.02196
- Published
- 2026-10-01
- Authors
- Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui
AI summary
Overview
Research area: Robot learning for humanoid loco-manipulation, combining behavioral foundation models (forward-backward representations), large language model driven reward design, and test-time search.
Technical level: Advanced. The paper assumes familiarity with reinforcement learning, reward shaping, successor-feature/forward-backward representations, and evolutionary strategies such as CMA-ES.
Scope: The paper proposes InterEvolve, a framework that adapts a single frozen humanoid whole-body controller to new contact-rich, multi-stage tasks by evolving staged reward programs at test time with an LLM agent and CMA-ES, retaining verified programs in a skill library.
What This Paper Is About
A humanoid controller trained on human-object interaction data can push, lift, and carry boxes, but it has no demonstration for a task like tipping a box onto another face, even though the motions required already exist inside it. The problem is that common task interfaces such as motion references, goal states, or skill labels require the user to specify desired motions or predefined behaviors, which is hard for nuanced contact-rich interactions. InterEvolve instead states a new task as an editable reward program and searches over those programs at test time, using execution feedback, without retraining the controller.
Key Contributions
- A framework for self-evolving humanoid loco-manipulation that shifts task-specific computation from training to test-time evolution: an LLM agent writes and adapts reward programs in context, and a fixed, reusable controller executes them for new tasks and scenes.
- An object-aware forward-backward (FB) behavioral foundation model that turns reward objectives into whole-body interaction, built by adding trainable object residuals to the frozen body networks of a pretrained body-only FB model (BFM-Zero) and training them on human-object interaction data.
- An evaluation of how reward design and accumulated experience affect reference tracking and goal-conditioned tasks, showing that success grows with evolution.
- Demonstration of unseen tasks and long-horizon compositions in simulation, plus fully autonomous deployment on a physical Unitree G1 from egocentric onboard perception.
Main Findings
- Object-aware pretraining supplies interaction competence. On the held-out large-box tracking benchmark, ULTRA keeps the lowest errors (E_h 15.68 cm, E_o 21.45 cm) with a success rate of 66 percent, and is a dedicated tracker trained for tracking alone. The body-only BFM-Zero loses the object on most clips (E_h 28.36, E_o 62.33, SR 8). InterEvolve without evolving reaches E_h 27.20, E_o 30.80, SR 60, and with an evolved tracking reward reaches E_h 19.80, E_o 24.92, SR 72, exceeding ULTRA's success rate while ULTRA retains lower errors.
- Reward design, not just weights, limits handcrafted rewards. Averaged over eight task families, an inaction baseline scores SR 0.0 and 49.1 earned tiers; a human-written program scores 8.0 SR and 65.2 earned; the agent's initial program scores 32.2 and 73.3; CMA-ES calibration alone raises these to 18.0/75.2 (human) and 34.6/82.5 (agent). InterEvolve reaches 86.5 SR and 95.6 earned, more than doubling the success of the best calibrated fixed program. Because both fixed programs are calibrated, the gap is attributed to structural changes.
- Staging and numerical calibration matter most. In the search-design ablation, the full system scores 86.5 SR and 95.6 earned at 2.1 GPU-h and 0.23M tokens. Forcing a single stage drops SR to 44.7; removing CMA-ES tuning drops it to 51.6; removing targeted edits (rewriting each program from scratch) gives 68.9; single-scenario evaluation gives 68.4; removing scene context gives 78.7.
- Success grows with test-time scaling, though not monotonically. Across rounds with controller weights fixed: initial 32.2 SR / 73.3 earned, round 1 75.0 / 91.7, round 2 76.8 / 92.6, round 3 73.6 / 93.1, round 4 (final) 86.5 / 95.6. Earned tiers rise monotonically while full success dips from round 2 to round 3, because success requires all criteria in the same rollout and fixing one criterion can briefly break another. LLM tokens grow from 0.03M to 0.23M, but simulation dominates the cost.
- Fixed programs fail where evolution succeeds. Calibrated fixed programs solve pushing to a mark but succeed in at most a quarter of episodes when carrying at chest height, tipping onto a new face, or pushing through a gate or around an obstacle; evolution lifts each of these families above 70 percent success. Kicking to a mark remains the hardest, since kicks are rare in the training data and need strong whole-body coordination from the reward.
- Reuse depends on covering the right contact modes. On three composite tasks with every condition searching for three rounds: the full library solves 8/10 relocation, 4/10 stacking, and 5/10 carry-place-kick episodes; no library gives 0/10, 0/10, and 1/10; a carry-only library gives 6/10, 1/10, and 1/10. Tokens exclude library construction (0.16M full, 0.13M none, 0.15M carry-only). Stored programs run without evolution almost never solve the task, so the gain comes from evolution integrating the library rather than direct retrieval.
- Long-horizon composition works in one take. The evolution agent arranges six scattered boxes into a ring, ordering the pushes itself; each leg runs a walking program to the box followed by the evolved pushing program, and one controller executes all twelve phases from a single reset. In the reported take, all six boxes end within 11 cm of their cells and no placed box is disturbed by later legs.
- Real-robot execution is autonomous but reported qualitatively. Programs evolved for kicking a box and pushing a box run on a physical Unitree G1 that detects the box with an egocentric camera and estimates its 6-D pose with FoundationPose. Quantitative real-world success rates are not reported in the content.
Methodology in Plain English
The system separates what a robot can do from what it should do. The "how to move" part is a fixed pretrained controller — an object-aware forward-backward behavioral foundation model. In a forward-backward model, a latent prompt indexes a policy, and a backward map projects any reward into that latent, so a new reward needs only a new prompt rather than a new policy. The authors take a body-only model and bolt trainable object residual branches onto its actor, forward map, and backward map, keeping each pretrained body branch frozen. The actor adds object features (position, orientation, linear and angular velocity, and distance-decayed vectors from body links to the nearest object surface) and outputs a mean action that is the frozen body prior plus an object-conditioned residual. To execute a stage reward, the model scores that reward over a fixed bank of body-object states sampled once from training replay, converts the scores into normalized weights, and projects the weighted sum of backward embeddings into the latent prompt at every control step.
The "what to do" part is a reward program: an ordered list of stages, each with reward code, a completion condition that reads live rollout context, and tunable constants such as weights, tolerances, kernel widths, and stage thresholds. Execution stays in a stage until its condition holds, then advances.
Evolution has two nested loops. In the outer loop, an LLM agent receives the task request, scene context, the fixed verifier criteria, the current best program with rollout feedback, and a skill library, then proposes new program structures; a validator discards programs that break code rules before any rollout. In the inner loop, CMA-ES tunes each valid proposal's constants within agent-declared bounds. Candidates are compared with the current best on the same scenarios, and the strongest is re-evaluated three times each against the current best and replaces it only if the improvement exceeds run-to-run evaluation noise. The next prompt reports outcomes per criterion and per stage. The budget is K revision rounds after the initial program; each round makes one agent call and spends 192 tuning rollouts per valid proposal plus 192 confirmation rollouts. Success is judged by a verifier whose criteria are instantiated once from the language task and scene and then stay fixed, so the agent cannot ease a task by rewriting its evaluator.
Pretraining data is human-object interaction from OMOMO and GRAB, retargeted to the Unitree G1 with rubber hands and to a G1 with Inspire hands. Following ULTRA, four box-like OMOMO objects are used: large box, plastic box, small box, and suitcase; for each object 50 clips are held out and the remaining 3,866 are used for training, with held-out large-box clips forming the tracking benchmark. DeepSeek-V4-Flash writes the reward programs, and all controllers run in Isaac Lab.
Why This Matters
Research impact. The paper argues for a different division of labor in humanoid intelligence: the controller learns how to move once, and task knowledge lives outside its weights as inspectable, reusable programs that a language model can read, edit, and test by execution. Reasoning and control improve on separate timescales — the controller with more interaction data, the agent with more test-time search and a growing program library — and neither requires retraining the other. It also shows that existing humanoid behavioral foundation models leave much of their competence untapped under human-designed rewards, which is a caution for how such models are evaluated.
Real-world applications (as implied by the demonstrated behaviors):
- Warehouse or logistics object handling, such as pushing boxes to marks, arranging boxes into configurations, and stacking one box on another.
- Contact-rich tool-free manipulation where the feet or other non-hand contacts are used, as in the feet-only kicking program.
- Long-horizon pick-and-place sequences, such as carrying a box, placing it, and then kicking it.
- Dexterous whole-body manipulation of small objects, shown qualitatively with Inspire hands on the G1, plus autonomous box pushing and kicking on a physical Unitree G1.
Industry relevance. The economics of the approach matter for deployment: each candidate reward costs a batch of rollouts on a fixed controller rather than a full RL run, and the reported cost for one complete evolution on one task family is 2.1 GPU-h and 0.23M LLM tokens on average over eight task families. Verified programs are stored as text and reused, so a fleet could in principle accumulate task knowledge as an editable document rather than as retrained weights.
Future Directions
- Cheaper evolution through faster simulation. The authors note that simulation dominates cost, with minutes per round for the language model against tens of minutes of rollouts, and suggest that a faster simulator such as mjlab in place of Isaac Lab could shorten evolution substantially.
- Quantitative real-robot evaluation. Real deployment on the G1 is shown for kicking and pushing boxes with onboard egocentric perception and FoundationPose, but quantitative real-world success rates are not reported, leaving physical robustness an open question.
- Broader skill-library coverage. Reuse tracked how well a library covered a task's contact modes: the carry-only library recovered most relocation but helped little on stacking or kicking, so what coverage is sufficient for general composition remains unresolved.
- Dexterous embodiments and novel contact modes. The Inspire-hand G1 results are qualitative, and the hardest family (kicking to a mark) remains the weakest, since such rare behaviors need strong whole-body coordination from the reward.
Target Audience
Robotics and embodied-AI researchers working on humanoid control, loco-manipulation, and behavioral foundation models; reinforcement learning researchers interested in reward design, LLM-driven reward search, and test-time compute; and engineers evaluating whether a fixed whole-body controller plus an evolving program library can substitute for task-specific policy retraining on real hardware.
Authors’ abstract
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new reward about the body or objects into loco-manipulation behavior at test time. Second, we specify tasks as reward programs: staged rewards with completion conditions and tunable constants. A large language model (LLM) agent revises the program structure in context, drawing on execution feedback and a skill library of verified programs, while a numerical optimizer tunes its constants. With every candidate verified across parallel simulation scenarios, the program explores new ways to induce, repurpose, and compose the controller's existing motor competence for the task at hand, and thus improves over iterations. Experiments show that human-designed rewards leave much of the FB model's loco-manipulation competence untapped, whereas the programs InterEvolve evolves release it, sometimes through novel strategies. It further produces behaviors for diverse tasks, complex scenes, and long-horizon compositions in simulation, and evolved skills run autonomously on a physical Unitree G1 from egocentric onboard perception.