Research
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental Adjustment Overview Research area: Computer Vision / robotics — specifically lifelong imitation learning (LIL) for robot manip

- arXiv
- 2603.10929
- Published
- 2026-03-11
- Authors
- Fanqi Yu, Matteo Tiezzi, Tommaso Apicella, Cigdem Beyan, Vittorio Murino
AI summary
Lifelong Imitation Learning with Multimodal Latent Replay and Incremental AdjustmentOverview
- Research area: Computer Vision / robotics — specifically lifelong imitation learning (LIL) for robot manipulation policies.
- Technical level: Intermediate. The paper assumes familiarity with imitation learning, experience replay, catastrophic forgetting, and Vision-Language Models such as CLIP, but its two core components (latent replay and an angular-margin regularizer) are conceptually simple.
- Scope in one sentence: The paper proposes and evaluates a rehearsal-based lifelong imitation learning framework that replays compact multimodal latent features instead of raw trajectories, plus a representation-regularization loss that keeps tasks separable, evaluated on three LIBERO manipulation benchmarks.
What This Paper Is About
Robots trained with standard imitation learning assume a fixed set of tasks, but real environments constantly introduce new objects, goals, and instructions. Lifelong imitation learning addresses this by letting a robot acquire new skills sequentially while retaining old ones, but the main obstacle is catastrophic forgetting: new-task training overwrites the representations that supported earlier tasks. This paper's goal is to reduce forgetting and improve transfer under realistic memory constraints, without needing a task identifier at test time and without fine-tuning large pretrained backbones.
Key Contributions
- Multimodal Latent Replay (MLR): a rehearsal mechanism that stores compact latent features — joint representations of vision, language, and robot state produced by the frozen encoders and modulation layer — together with their control commands, rather than raw images and trajectories. This reduces the memory footprint of previously learned tasks relative to storing raw data.
- Incremental Feature Adjustment (IFA): a representation-level regularization loss computed on angular distances that repels the current task's global latent representation from previously learned task references while keeping it attracted to its own reference, preserving inter-task distinctiveness.
- Adaptive margin: the IFA margin is not fixed; it is scaled by the distance between task references (controlled by a factor alpha), so the loss strength adapts to the semantic similarity between tasks, removing the need for manual per-dataset tuning.
- An architecture that requires no PEFT and no distillation: all encoders remain frozen during the lifelong stage — only the temporal decoder and policy head are updated — yet the method reports state-of-the-art performance on three LIBERO benchmarks in a task-ID agnostic setting.
Main Findings
- New state of the art on LIBERO: The full method (MLR + IFA) achieves the highest FWT and AUC and among the lowest NBT across LIBERO-OBJECT, LIBERO-GOAL, and LIBERO-50. The abstract reports 10–17 point gains in AUC and up to 65% less forgetting compared to previous leading methods.
- Latent replay alone already beats prior methods: MLR alone reaches an AUC of 77.6 ± 3.0 on LIBERO-OBJECT, 74.6 ± 2.7 on LIBERO-GOAL, and 54.7 ± 2.4 on LIBERO-50, above LOTUS (65.0, 56.0, 45.0) and M2Distill (69.0, 57.0; not available on LIBERO-50).
- IFA improves every metric on every benchmark: On LIBERO-OBJECT, AUC rises from 77.6 ± 3.0 (MLR) to 79.4 ± 1.5 and NBT falls from 12.3 ± 2.4 to 11.4 ± 5.6. On LIBERO-GOAL, AUC rises from 74.6 ± 2.7 to 77.2 ± 1.8 and NBT falls from 10.0 ± 7.5 to 6.9 ± 0.9. On LIBERO-50, AUC rises from 54.7 ± 2.4 to 56.1 ± 1.8 and NBT falls from 20.6 ± 9.0 to 8.6 ± 6.2.
- Reduced forgetting versus the closest competitors: On LIBERO-GOAL, the paper contrasts its NBT of 6.9 against ISCIL's 19.4 and M2Distill's 20.0, while AUC improves over ISCIL's 60.5.
- Language + agent-view is the best modality pairing for task-pair selection: In the modality ablation, "Lan + AV" achieves the best LIBERO-OBJECT AUC (79.4 ± 1.5) and LIBERO-GOAL AUC (77.2 ± 1.8) among the tested single modalities and pairings.
- Top 50% of most similar task pairs is the best selection proportion: For LIBERO-OBJECT, 50% gives FWT 84.6 ± 1.9, NBT 11.4 ± 5.6, AUC 79.4 ± 1.5, versus AUC 77.6 ± 3.0 at 33.3% and 75.8 ± 5.3 at 66.6%. On LIBERO-GOAL, 33.3% and 50% tie (80.0 / 6.9 / 77.2), while 66.6% is worse on AUC (75.6 ± 8.7).
- Language embeddings are better references than averaged global representations: Language references give LIBERO-OBJECT AUC 79.4 ± 1.5 versus 75.7 ± 1.1 for "mean global," and LIBERO-GOAL AUC 77.2 ± 1.8 versus 70.5 ± 6.0.
- Larger buffers help, and even small ones work well: On LIBERO-OBJECT, a 0.50 storage probability corresponds to 188MB and AUC 79.4 ± 1.5; 0.20 gives 75.3MB and 77.7 ± 10.0; 0.10 gives 37.6MB and 76.6 ± 14.1. On LIBERO-GOAL the equivalents are 121.2MB / 77.2 ± 1.8, 60.6MB / 77.4 ± 4.6, and 30.3MB / 66.2 ± 6.8.
- Angle-based loss beats cosine-distance-based loss: The angle-based formulation outperforms the cosine variant on all metrics, and the cosine variant still surpasses prior SOTA and improves on MLR-only, though with notably higher standard deviation.
- IFA produces more separated task clusters: UMAP visualizations show that without IFA, embeddings of different tasks are intertwined, while with IFA they form compact, distinct clusters around task-specific references. Quantitatively, on LIBERO-OBJECT the cosine similarity between task 8 and task 7 drops from 0.1374 to 0.1218, and on LIBERO-GOAL the similarity between task 8 and task 6 drops from 0.2191 to 0.1977 — in both cases, the only pair to which IFA was applied.
- Failure mode acknowledged: Tasks that are already highly similar or overlapping may remain so, because IFA targets high-similarity cases with an adaptive angular margin that preserves small distinctions which cosine similarity compresses.
Methodology in Plain English
The pipeline has two stages. First, a multi-task pretraining stage trains the whole architecture offline on a disjoint set of tasks. The architecture encodes three modalities: images via a frozen CLIP image encoder, language instructions via a frozen CLIP text encoder, and robot proprioceptive state via a two-layer MLP. A FiLM layer modulates the visual and state features using language features, producing task-conditioned representations. These are concatenated over the temporal window into a multimodal sequence, passed through a GPT-2 decoder as the temporal decoder, and finally through a policy head that predicts the next action. Training uses a standard behavior cloning loss.
Second, a lifelong learning stage presents new tasks one after another. Only the temporal decoder and policy head remain trainable; everything else stays frozen. Rather than storing pixels, the method stores the already-computed multimodal latents (the concatenated encoder outputs) along with their actions in a bounded replay buffer, with capacity balanced across previously seen tasks and sized to roughly five demonstrations per task. When a new task arrives, the model trains on the union of the new task's data and replayed latents.
On top of behavior cloning, the IFA loss is added. For each selected pair consisting of one old and one newly introduced task, the loss requires the current task's global latent representation to be closer to its own task reference than to the old task's reference by at least a margin. The reference used is the language embedding of the task, because it is stable and fixed. The margin itself is not a constant: it equals a scaling factor alpha times the angular distance between the two task references, so it grows when the tasks are dissimilar in reference space. Distance is measured as the angle between vectors (arccos of cosine similarity), which the paper says magnifies differences between close references compared to plain cosine distance. Task pairs are selected by computing cosine similarity in the language and agent-view modalities across all timesteps, ranking pairs, and keeping only those that fall in the top 50% of most similar pairs in both modalities and that pair a new task with a previously learned one.
Training settings: AdamW with a learning rate initialized to 10^-4 and a linear scheduler, batch size 10 for both stages, 100 epochs, lambda_IFA fixed at 0.1, and alpha set to 0.3, 0.7, and 0.1 for LIBERO-OBJECT, LIBERO-GOAL, and LIBERO-50 respectively. Experiments run on a single NVIDIA A100 GPU, repeated over three random seeds with 20 trials per seed.
Why This Matters
Impact on research. The paper shows that strong lifelong imitation learning can be achieved with entirely frozen pretrained backbones, without parameter-efficient fine-tuning, without knowledge distillation, and without task identifiers at test time. It also decouples memory cost from raw-data storage by replaying latents, and it demonstrates that an adaptive angular-margin regularizer can substitute for more elaborate machinery. This gives the LIL community a simpler, more memory-efficient baseline and a new SOTA to compare against on LIBERO.
Real-world applications (bullets):
- Household service robots that must pick up newly introduced kitchen tools, navigate rearranged furniture, or handle tasks they were never pretrained on.
- Warehouse or logistics robots whose object sets and pick-and-place goals change over deployment time.
- Assistive or care robots whose environment and user instructions evolve, requiring continuous skill addition without retraining from scratch.
- Any deployed manipulation system with constrained onboard memory, since latent replay reduces buffer footprint compared to storing raw image trajectories.
Industry relevance. The method's reliance on frozen CLIP and GPT-2 components means practitioners can adopt it without expensive backbone fine-tuning infrastructure, and the bounded, task-balanced buffer makes memory budgeting predictable. The reported buffer sizes (for example, 37.6MB to 188MB on LIBERO-OBJECT depending on storage probability) make the trade-off between memory and performance explicit, which matters for edge deployment. The code is released at https://github.com/yfqi/lifelong_mlr_ifa.
Future Directions
- Handling highly similar tasks. The authors note that very similar tasks may still remain highly similar or overlapping even with IFA. A method that can separate near-duplicate task references, or that explicitly detects when separation is impossible, remains open.
- Extending beyond the LIBERO simulation suite. All reported experiments are on LIBERO-OBJECT, LIBERO-GOAL, and LIBERO-50; the paper does not report results on physical robots, so transfer to real hardware and real sensor noise is untested.
- Scaling the task sequence. LIBERO-50 uses five stages of five new tasks each, and M2Distill results are not available (marked "NA") for that suite, leaving open how the method behaves over much longer task sequences and how the buffer balancing strategy scales.
- Reducing the storage budget further. Since performance degrades as the storage probability drops from 0.50 to 0.10, an open question is whether smarter selection of which latents to store could maintain performance at much smaller buffer sizes.
Target Audience
This paper is most useful to robotics and embodied-AI researchers working on continual learning, imitation learning, or policy learning from demonstrations; to machine learning engineers building deployed manipulation systems that must adapt to new tasks without full retraining; and to graduate students who want a concrete, well-ablated example of combining rehearsal-based replay with a representation-level regularizer. Readers should be comfortable with behavior cloning, experience replay, CLIP-style vision-language encoders, and lifelong learning metrics such as FWT, NBT, and AUC.
Authors’ abstract
We introduce a lifelong imitation learning framework that enables continual policy refinement across sequential tasks under realistic memory and data constraints. Our approach departs from conventional experience replay by operating entirely in a multimodal latent space, where compact representations of visual, linguistic, and robot's state information are stored and reused to support future learning. To further stabilize adaptation, we introduce an incremental feature adjustment mechanism that regularizes the evolution of task embeddings through an angular margin constraint, preserving inter-task distinctiveness. Our method establishes a new state of the art in the LIBERO benchmarks, achieving 10-17 point gains in AUC and up to 65% less forgetting compared to previous leading methods. Ablation studies confirm the effectiveness of each component, showing consistent gains over alternative strategies. The code is available at: https://github.com/yfqi/lifelong_mlr_ifa.