Research
In-Context Learning for Robots: Methods and Applications
Overview Research area: Robotics, specifically in-context learning (ICL) for robots — how demonstrations, instructions, corrections and interaction history can redirect a robot's behavior while its ne

- arXiv
- 2609.36012
- Published
- 2026-09-28
- Authors
- Haojian Huang, Zexi Li, Junhao Guo, Yehang Zhang, Wenxuan Peng, Bohan Zhou, Weilin Ruan, Leyi Wu, Chenxu Wang, Jianchong Su, Binghui Xie, Wosong Chen, Yingjie Xu, Tianhao Zhou, Suzeyu Chen, Pukun Zhao, Jiaqi He, Xinyi Li, Runze Li, Peiran Dong, Shaoxiang Dang, Jing Huang, Yingbing Chen, Yifan Chang, Tianyi Zhang, Shiyuan Deng, Haozhi Wang, Yangkai Wei, Wenqian Li, Han Yang, Kaiwen Zhou, Huaping Liu, James Cheng, Rui Shao, Donglin Wang, Yaochu Jin, Jianye Hao, Ying-Cong Chen, Yinchuan Li
AI summary
Overview
Research area: Robotics, specifically in-context learning (ICL) for robots — how demonstrations, instructions, corrections and interaction history can redirect a robot's behavior while its neural parameters stay fixed during deployment.
Technical level: Advanced. This is a literature review that assumes familiarity with imitation learning, reinforcement learning, vision-language-action models, world models, and manipulation/navigation benchmarks.
Scope: The paper organizes the literature on ICL for robots into four method families and six "learning horizons," and connects method design to how transfer and evaluation are judged.
What This Paper Is About
General-purpose robots must infer what a new task requires and turn that understanding into physical action, but broader motor competence alone cannot resolve everything: the current scene may leave several procedures feasible, hide an earlier event, or reveal little about an unfamiliar material's response. In-context learning for robots addresses these information gaps by using demonstrations and interaction to direct existing competence, with neural parameters held fixed during deployment — the alternative, fine-tuning, bakes new evidence into weights instead. The paper's organizing question is how new evidence resolves what existing competence leaves undetermined, and how that resolution survives physical execution.
Key Contributions
- A four-family taxonomy. The review distinguishes context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution, and identifies correspondence and memory mechanisms that connect methods across these families.
- An account of how training relationships establish context use — and where transfer breaks. The paper explains how object substitution, unfamiliar environments, and execution conditions limit transfer, and it separates broader motor competence from a broader ability to learn through teaching.
- A synthesis of reported comparisons and evaluation controls. The review separates three things that are often conflated: responsiveness to teaching, physical transfer, and benefits from retained experience. This motivates compositional task acquisition and improved teachability.
- A placement of the review against adjacent surveys. The introduction surveys prior reviews of learning from demonstration, in-context RL, human-video learning, manipulation ICL, VLM-based VLA models, cross-embodiment adaptation, predictive control, and skill acquisition; a table in the Appendix compares related reviews, including recent world-action taxonomies.
Main Findings
- Four interfaces connect context to execution. Context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution each place the burden of "what to do" in a different intermediate, and each carries different transfer assumptions.
- Six learning horizons describe where learning changes the system. S1 encodes control explicitly (task logic, control laws, planning models); S2 learns reusable policies through imitation and reinforcement; S3 adapts to interaction outcomes using history; S4 infers new task requirements from supplied teaching. S5 and S6 are research objectives: S5 is physical recursive self-improvement, in which experience improves how the next task is learned; S6 extends this through validated knowledge exchange across embodiments. S3 and S4 both operate with deployed neural parameters fixed.
- Fine-tuning and ICL differ by where task information is stored. Fine-tuning uses task data to update pretrained parameters into task-adapted parameters; the illustrated ICL case supplies a support demonstration as context while keeping parameters fixed. Both use observation–action history to predict the next action block. Storage location determines acquisition cost, persistence, and reset behavior.
- Context use depends on relationships learned in training. The paper cites GPT-3's few-shot performance after autoregressive pretraining and controlled studies identifying distributional conditions that promote using examples. For robots, the critical learned relationship links earlier teaching or interaction to the later action it changes — and retaining a correction can extend that relationship across attempts, provided its conditions remain valid.
- Cross-object transfer is a core test. Teaching must remain useful after replacing either or both interacting objects, so the task relation must survive while its physical realization changes. Placing an object at the correct destination can still violate a required handle grasp, so both intended effect and prescribed order/contact must be traced through execution.
- Collective knowledge exchange rests on three complementary foundations. Shared progress and skill representations preserve what should be accomplished; pooled policy learning supplies competence across sensing and action spaces; morphology-aware transfer changes the physical realization. The premise is that an execution contains knowledge another body can use even when the original motion cannot be copied.
- Retention is conditional, not automatic. A retained correction can guide both interpretation and execution, possibly sharing a model and input modality — but its validity depends on conditions that may change.
- No quantitative results are reported in the available content. The paper is a review; the text supplied does not include benchmark scores, dataset sizes, or numerical comparisons between methods.
Methodology in Plain English
This is a literature review, not an experimental study. The authors begin from a single organizing question — how new evidence resolves what existing competence leaves undetermined, and how that resolution survives physical execution — and use it to sort the field. They group methods into four families by which intermediate object carries the task information (policy context, geometric correspondence, world-model predictions, or executable skills/agents), then examine what each family assumes about correspondence, training, and memory. They place the six learning horizons alongside these families as a way of distinguishing programmed competence, learned competence, interaction-based adaptation, teaching-based adaptation, improvement in learning itself, and knowledge exchange across robots. Notation for contextual evidence, execution intermediates, and retained state is collected in one table so the alternatives can be compared through a common policy interface. The argument is laid out as a sequence: why context is needed, how context changes action, how that dependence is learned, where its physical limits appear, how to evaluate it, and what research agenda follows.
Why This Matters
Research impact. The review's value is organizational and diagnostic. By asking which intermediate carries the task information, it makes the transfer assumptions of different method families comparable, and by separating context dependence, physical transfer, and retained-experience benefits, it gives evaluation a vocabulary for testing what a system actually learned from context. The distinction between broader motor competence and broader teachability reframes progress in generalist robot policies.
Real-world applications (as discussed or implied in the paper):
- Assembly and packing. Demonstrations specify fold or assembly order, history identifies completed placements, and contact evidence reveals a tighter fit or misalignment.
- Navigation. Earlier visits reveal locations outside the current view, so a robot can act on places it cannot currently see.
- Cross-object manipulation. Teaching must survive substitution of the objects involved, which is the practical case in unstructured settings.
- Cross-embodiment skill sharing. Robots of different bodies exchanging demonstrations and corrections, where shared progress and skill representations preserve what should be accomplished even when motion cannot be copied.
- Automated practice and reward/experiment design, cited as mechanisms supporting improvement in how subsequent tasks are learned, along with retained guidance from earlier corrections.
Industry relevance. The paper situates ICL against the foundation-model stack for robotics: broad pretraining from multi-task and multi-embodiment collections, and generalist policies that teachable context is meant to steer. For companies deploying general-purpose robots, the review's core claim — that the same visible scene can require different actions, and that context resolves that ambiguity using actions the robot can already execute — bears directly on whether a deployed system can be redirected without another task-specific parameter update.
Future Directions
- Compositional task acquisition and faithful transfer. The agenda connects learning new tasks by composing known relations with preserving taught requirements as objects, environments, and execution conditions change.
- Physical recursive self-improvement. Making experience improve the ability to learn subsequent tasks is framed as an objective (S5), realizable through retained context, executable programs, or neural updates — and the paper notes their differing evaluation.
- Collective knowledge evolution. Whether exchanges of demonstrations and corrections actually improve a group of robots' subsequent learning (S6) remains an open research objective rather than a demonstrated result.
- Evaluation that distinguishes mechanisms. Better controls are needed to separate context dependence from transfer and from retained-experience benefit, and to test the validity of context when environmental conditions are non-stationary.
- Establishing when retained guidance stays valid. Retained corrections extend teaching across attempts only while their conditions hold, which raises the question of how to detect and handle conditions that have changed.
Target Audience
Robotics researchers working on imitation learning, in-context learning, manipulation, and navigation; engineers building vision-language-action and generalist robot policies who need to know when a deployed system can be redirected by teaching rather than retraining; and evaluation-focused researchers looking for a framework that distinguishes teaching responsiveness, physical transfer, and retained-experience benefits. Readers without background in robot learning will find the conceptual structure accessible, but the surrounding citations and terminology assume an advanced audience.
Authors’ abstract
General-purpose robots must infer what a new task requires and translate that understanding into appropriate physical action. In-context learning (ICL) for robots supports this process by using demonstrations and interaction to direct existing competence with neural parameters held fixed during deployment. We organize this literature review around the interfaces connecting contextual evidence to execution, distinguishing four families: context-conditioned policies, geometric demonstration transfer, world-model-based control, and skill- and agent-based execution. Comparing these interfaces clarifies their transfer assumptions and the roles of training, correspondence, and memory in making context useful. Across manipulation and navigation, we examine how these mechanisms preserve taught requirements as objects, environments, and execution conditions change. This analysis links method design to evaluation practices that distinguish responsiveness to teaching, physical transfer, and benefits from retained experience. The resulting agenda connects compositional task acquisition and faithful transfer with physical recursive self-improvement, in which experience improves the ability to learn subsequent tasks.