Research
ARM-FM: Automated Reward Machines via Foundation Models for Compositional Reinforcement Learning
Overview Research area: Reinforcement learning (reward specification and reward machines), foundation models for sequential decision-making, and compositional/multi-task RL. Technical level: Advanced.

- arXiv
- 2510.14176
- Published
- 2025-10-16
- Authors
- Roger Creus Castanyer, Faisal Mohamed, Pablo Samuel Castro, Cyrus Neary, Glen Berseth
AI summary
Overview
Research area: Reinforcement learning (reward specification and reward machines), foundation models for sequential decision-making, and compositional/multi-task RL.
Technical level: Advanced. The paper assumes familiarity with Markov Decision Processes, automata-based reward machines, and deep RL algorithms (DQN, PPO, SAC, Rainbow DQN).
Scope: The paper introduces ARM-FM, a framework that prompts foundation models to automatically construct language-aligned reward machines from natural-language and visual task descriptions, then conditions a single RL policy on the language embeddings of the machine's states to enable dense rewards, skill reuse, and zero-shot generalization.
What This Paper Is About
Reinforcement learning agents are highly sensitive to how their reward function is specified: sparse rewards give too weak a learning signal, while hand-crafted dense rewards invite reward hacking. Reward machines offer a principled, structured alternative but have historically required manual, expert-driven design, which limits them to task-specific applications. The paper's goal is to automate the construction of reward machines using foundation models and to turn their states into a shared, language-grounded skill space so that one policy can learn and transfer across many tasks.
Key Contributions
- Automated LARM generation. A framework that generates complete task specifications directly from natural language (and visual observation) using foundation models, producing language-aligned reward machines (LARMs) that include the automaton structure, executable labeling functions, and natural-language instructions for each subtask.
- A shared, language-grounded skill space. A method that exploits the language-aligned nature of the resulting automata to let policies share knowledge across related subtasks, enabling experience reuse and policy transfer across related tasks.
- Empirical validation across domains. Experiments showing the framework solves long-horizon, sparse-reward tasks in grid worlds, a 3D procedurally generated Minecraft-based world, and continuous-control robotic manipulation, while supporting multi-task training and zero-shot generalization.
- Analysis of FM-generated components. An LLM-as-judge evaluation over 1,000 sampled tasks showing a scaling trend in the reliability of generated reward machines and labeling code, plus a PCA analysis showing semantic clustering in the state-instruction embedding space.
Main Findings
-
Sparse-reward grid tasks are solved where baselines fail. On MiniGrid-DoorKey at increasing grid sizes, both on fixed maps and procedurally generated layouts, the DQN+RM agent consistently outperforms DQN+ICM, ReAct, and unmodified DQN. On the harder UnlockToUnlock, BlockedUnlockPickup, and KeyCorridor tasks, the authors state their agent is the only one of the compared methods to solve all three and reach near-perfect reward, while the baselines show no learning.
-
Dense structured rewards transfer to a 3D procedural world. In Craftium, where the agent must gather wood, stone, and iron in order to mine a diamond and only a sparse reward is given at the end, PPO augmented with a generated LARM consistently completes the entire task sequence, whereas the baseline PPO agent makes minimal progress.
-
Robotic manipulation benefits from automatic reward engineering. On five Meta-World tasks (Assembly, Bin-Picking, Pick-Place, Shelf-Place, Stick-Push) using SAC, the method achieves higher success rates than learning from the sparse reward alone. Appendix results also report that with careful hyperparameter tuning the RM-augmented agent can reach a high success rate, and that combining the reward machine with an RND exploration bonus yields better overall performance in most environments than the main Meta-World results.
-
Both structured rewards and language embeddings are necessary for multi-task robustness. In the XLand-MiniGrid ablation, a single Rainbow DQN agent is trained on an increasing number of simultaneous tasks (1, 3, 5, and 10). The baseline fails to generalize as tasks increase; using only state embeddings gives a weak learning signal that degrades quickly; using only LARM rewards enables multi-task learning but the policy struggles because it is unaware of the active sub-goal; only the full method maintains high success as the number of tasks grows.
-
Zero-shot generalization to a novel composite task. An agent trained on tasks A and B with their reward machines solves a new, unseen task C with a novel FM-generated LARM, without any fine-tuning, when the sub-goals of C are semantically familiar from training (for example, "Pick up a blue key").
-
Bigger foundation models generate better reward machines. Across 1,000 sampled XLand-MiniGrid tasks judged by Qwen3-30B-A3B-Instruct-2507, the LLM-as-judge evaluation shows a clear scaling trend, with larger models such as Qwen3-32B significantly more capable of producing fully correct task specifications. Mistral-Small is noted as more adept at generating a valid RM structure than correct labeling code.
-
State instructions form a semantically coherent embedding space. PCA of embeddings of state instructions from the 1,000 generated tasks shows distinct clusters, with start, middle, and end states occupying different regions and semantically similar instructions from different tasks clustering together.
-
Theoretical grounding. The paper states that well-designed LARMs yield theoretical guarantees ensuring the generated reward structure preserves the optimal policy of the original sparse task, with details in Appendix A.5.
Methodology in Plain English
The authors treat a reward machine as a small finite-state automaton whose states correspond to subtasks. Given a high-level natural-language prompt plus a visual observation of the environment, a foundation model writes out three things: the automaton's structure, executable Python labeling functions that detect when symbolic events occur in the environment, and a natural-language description of each state. The specification is refined over N rounds in a self-improvement loop using paired generator and critic foundation models, with an optional human approval or correction step.
Because each automaton state now has a natural-language description, the authors embed those descriptions into vectors. During training, the RL agent's policy takes the environment observation together with the embedding of the current machine state, so the agent always knows which sub-goal is active. The environment observation and the machine reward are combined into an augmented learning problem: the agent receives the sum of the original environment reward and the reward machine's reward, and the labeling functions drive state transitions.
This design means related subtasks ("pick up a blue key" versus "pick up a red key") sit close together in embedding space, so a single policy can reuse behavior across tasks and across different reward machines. The authors evaluate this on MiniGrid and BabyAI (with DQN), Craftium (with PPO), Meta-World (with SAC), and XLand-MiniGrid (with Rainbow DQN for multi-task experiments). All LARM components were generated with GPT-4o except the 1,000 XLand-MiniGrid reward machines used in the ablation study, which came from various open-source foundation models. Reported results are averaged over 3 independent random seeds, with shaded regions and error bars indicating one standard deviation.
Why This Matters
Impact on research. The work connects two research threads that are usually separate: the formal, verifiable structure of reward machines and the semantic reasoning of foundation models. It suggests a principled path toward hierarchical, interpretable RL where task specifications are human-readable, inspectable, and modifiable rather than opaque learned reward models. It also provides evidence that language-conditioned policies can function as a compositional library of reusable skills within the automaton formalism.
Real-world applications (from the domains the paper evaluates):
- Robotic manipulation and assembly, where the paper shows automatically generated dense rewards reduce the need for hand-engineering low-level signals such as joint angles.
- Resource-gathering and crafting-style open worlds, demonstrated in the Minecraft-based Craftium environment.
- Instructable agents in grid-like navigation and planning domains, such as the nested key-and-door tasks in MiniGrid and BabyAI.
- Multi-task training systems that must adapt to newly sampled tasks, as tested in XLand-MiniGrid.
Industry relevance. Reward engineering is a well-known cost driver in applied RL, and this framework targets that cost directly by turning a natural-language description into a structured reward signal. Because the generated reward machines are interpretable and language-based, they also offer an interface for human oversight and refinement of an agent's objectives. The paper acknowledges funding from NSERC, Google Research, and CIFAR, and compute support from the Digital Research Alliance of Canada, Mila IDT, and NVidia.
Future Directions
- Reducing or removing the reliance on human verification during reward machine generation, for example by exploiting the automaton structure to enable automated self-correction through formal verification.
- Extending zero-shot generalization beyond novel task compositions within the same domain, since the paper explicitly scopes its zero-shot claims to that setting rather than cross-domain transfer.
- Broadening the labeling functions beyond Python code, since the framework is described as general and able to support any boolean predicate such as formal logic or queries to other foundation models.
- Further investigation of model-specific strengths in generation, given the finding that some models produce valid automaton structures more reliably than they produce correct labeling code.
Target Audience
Researchers and practitioners working on reinforcement learning who struggle with reward specification, particularly those interested in reward machines, hierarchical and compositional RL, multi-task learning, and zero-shot generalization. It is also relevant to readers studying the interface between foundation models and decision-making agents, and to engineers who want an interpretable, language-driven way to specify agent objectives. A background in RL and automata-based task specification is needed to follow the formal definitions and experimental setups in detail.
Authors’ abstract
Reinforcement learning (RL) algorithms are highly sensitive to reward function specification, which remains a central challenge limiting their broad applicability. We present ARM-FM: Automated Reward Machines via Foundation Models, a framework for automated, compositional reward design in RL that leverages the high-level reasoning capabilities of foundation models (FMs). Reward machines (RMs) -- an automata-based formalism for reward specification -- are used as the mechanism for RL objective specification, and are automatically constructed via the use of FMs. The structured formalism of RMs yields effective task decompositions, while the use of FMs enables objective specifications in natural language. Concretely, we (i) use FMs to automatically generate RMs from natural language specifications; (ii) associate language embeddings with each RM automata-state to enable generalization across tasks; and (iii) provide empirical evidence of ARM-FM's effectiveness in a diverse suite of challenging environments, including evidence of zero-shot generalization.