Research
Arm-wise Compositional Generalization in Dual-Arm Vision-Language-Action Models
Overview Research area: Robotics — dual-arm (bimanual) vision-language-action (VLA) manipulation policies, specifically compositional generalization over arm-level skills. Technical level: Advanced. T

- arXiv
- 2610.06184
- Published
- 2026-10-05
- Authors
- Zaibin Zhang, Binghao Ran, Yuhan Wu, Zhongbo Zhang, Yifan Wang, Junwei Jiang, Junlan Xiao, Wangcheng Shi, Li Kang, Yiran Qin, Zhenfei Yin, Lijun Wang, Huchuan Lu
AI summary
Overview
- Research area: Robotics — dual-arm (bimanual) vision-language-action (VLA) manipulation policies, specifically compositional generalization over arm-level skills.
- Technical level: Advanced. The paper assumes familiarity with VLA backbones, flow-matching action heads, LoRA adapters, and attention masking, though the central question is stated plainly.
- Scope: A benchmark (ACG-Bench) plus a controlled architectural study on a shared π0.5 backbone that tests whether familiar per-arm skills can be recombined under unseen ordering, timing, and cross-task requirements, with a combined design called AE-VLA.
What This Paper Is About
Dual-arm VLA models can learn manipulation tasks, but success on training tasks does not show whether their component arm-level skills can be reused in new combinations. The authors define Arm-wise Compositional Generalization (ACG) as reusing familiar arm-level skills under coordination requirements never seen in training — different orderings, synchronized completions, combinations of the two, and pairings of skills drawn from different tasks. The paper builds a benchmark to measure this and runs a controlled design study to identify which training and architectural choices support it, rather than proposing a single monolithic new system.
Key Contributions
- ACG-Bench: a benchmark for systematic evaluation of dual-arm skill recomposition, containing 23 task–condition pairs across 8 task families — 6 in-domain conditions and 17 unseen compositions spanning Reorder, Sync, Sync+Reorder, and Cross-task Combination. Success requires both the task goal and compliance with physical milestones, order constraints, and timing windows.
- A unified empirical study: comparisons of representative data-augmentation strategies (MA-VLA's Arm Shuffle) and architectural strategies (arm-token grouping, SkillLoRA, arm-wise attention), all sharing a common π0.5 backbone, the same pooled source dataset, and one execution protocol and scheduler.
- Design insight on complementarity: evidence that skill-conditioned parameters and arm-wise attention interact strongly, quantified through a 2×2 comparison of SkillLoRA and AWA on top of token grouping.
- AE-VLA: the combined configuration, validated in simulation (21.53% generalization CCSR) and on physical SO101 robots (39.00% mean success across five unseen conditions).
Main Findings
- Large in-domain versus compositional gap: Single π0.5 falls from 37.00% in-domain CCSR to 2.94% on the 17-condition generalization set. MA-VLA reaches 3.06% and two independently controlled π0.5 policies (Dual π0.5) reach 5.53%. Arm-role augmentation and independent control both provide only limited help.
- AE-VLA leads generalization: the combined design reaches 21.53% generalization CCSR — a gain of 18.59 points over Single π0.5 and 16.00 points over Dual π0.5.
- Group-level differences: AE-VLA reaches 26.50% on Reorder, 22.00% on Sync+Reorder, and 40.00% on Cross-task Combination, while pure Sync remains difficult at 8.67%.
- Progress does not equal success: AE-VLA and Dual π0.5 complete similar fractions of required milestones on the generalization set (44.25% and 40.64%), yet their full success rates differ widely. Single and MA-VLA show low progress (10.94% and 9.83%). On synchronization tasks, Dual π0.5 can grasp objects but lacks the waiting and coordination, producing arm-wise collisions and failure.
- Individual architectural changes give modest gains: Token Group raises generalization CCSR from 2.94% to 3.82%; adding SkillLoRA alone reaches 6.00%; adding AWA alone reaches 5.82%.
- Strong interaction between SkillLoRA and AWA: the AWA gain is 2.00 points without SkillLoRA and 15.53 points with it, an interaction of 13.53 points. Jointly they reach 21.53%.
- In-domain cost of AWA: AWA lowers Token Group from 45.00% to 25.00% in-domain CCSR, and Token Group + SkillLoRA from 41.33% to 27.17%. AE-VLA's in-domain success of 27.17% is below all three baselines, so improved generalization comes with reduced in-domain performance.
- Baselines still win some conditions: Dual π0.5 is stronger on Object in Cabinet/In-domain (59.00 vs 21.00) and Burger & Fries/Sync (6.00 vs 1.00). Single π0.5 is best on Push Cubes/In-domain (50.00) and MA-VLA is best on Stack Cubes/In-domain (28.00).
- Large gains on specific conditions: AE-VLA reaches 79.00% on Stack Bowls/Reorder versus 4.00% for Dual, and 54.00% on Cube in Bowl/Cross-task versus 0.00% for Dual.
- Physical robot results: on two 6-DoF SO101 arms with one global and two wrist cameras, AE-VLA reaches 39.00% mean success over five unseen conditions, versus 2.00% for Single π0.5, 6.00% for MA-VLA, and 10.00% for Dual π0.5. It leads on four of five unseen conditions, including Stack Bowls/Sync (12/20 versus 0/20 for every baseline); Push Cubes/Sync is the exception (Dual 6/20, AE-VLA 2/20). In-domain, AE-VLA reaches 60.00%, below Dual's 70.00%.
Methodology in Plain English
The authors build a benchmark in ManiSkill using the task and expert-program generation pipeline of RoboTwin 2.0. Six source task families (Stack Bowls, Stack Cubes, Push Cubes, Burger & Fries, Can in Basket, Object in Cabinet) are generated and validated by checking agents that repair failed grasps, collisions, and failed task checks; accepted programs collect 100 clean expert demonstrations per source task, pooled into one multitask training set of 600 demonstrations. Two families, Cube in Bowl and Bowl in Cabinet, are held out as cross-task tests.
Each task is described as an event graph: nodes are skill executions with a specified arm, object, and target, with order constraints on which events precede others and synchronization constraints requiring two completion milestones to fall within a task-specific time window. A test condition counts as ACG when its skill types appear in training but their joint event graph does not.
Crucially, a common rule-based scheduler supplies per-arm atomic prompts to every method at evaluation, and shared state checks update those prompts and record milestone completion times. This isolates execution of a supplied composition from planning, so the study does not measure whether a model can infer a plan.
On top of a shared pretrained π0.5 backbone, three architectural changes are tested. Arm-token grouping replaces the H=50 joint action tokens with two groups of H=50 tokens, one per arm, sharing one action expert and output projection. SkillLoRA adds a bank of K=10 low-rank updates indexed by atomic skill type (rank k=4, α=16), applied to attention projections and feed-forward layers in all 18 action-expert blocks and to the action input/output projections, with a separate two-layer router per arm selecting the adapter from prefix features. Arm-wise attention (AWA) restricts which keys each query group can attend to — global tokens read both local prefixes, each arm's prefix reads only the global tokens and its own prefix, and each action token reads only its own arm's earlier action tokens, with prefix tokens never attending to action tokens.
The four grouped variants form a 2×2 comparison of SkillLoRA and AWA. Training uses flow matching with an action loss plus 0.1 times the router cross-entropy loss for SkillLoRA variants, jointly fine-tuning the backbone, action expert, adapters, and routers for 30,000 steps with batch size 32 on two NVIDIA A800 GPUs. Evaluation uses 100 episodes per condition with seeds 1000–1099, giving 2,300 episodes across the 23 conditions. The main metric is constraint-compliant success rate (CCSR), which counts only episodes that reach the goal and satisfy the required milestones, order, and timing; normalized progress measures the fraction of required milestones completed. For Sync, the allowable event-time gap is 20 s for Can in Basket, 15 s for Object in Cabinet, and 45 s for all other tasks. Arm-wise collisions zero the success score while retaining measured progress. TwinVLA is omitted from the quantitative comparison because its pretrained weights differ from π0.5.
Why This Matters
Impact on research. The paper reframes dual-arm generalization as a compositional problem over arm-level skills rather than a matter of learning more coordination patterns, and it supplies a shared protocol in which backbone, data, prompts, and metrics are held fixed. This makes architectural claims about arm-role augmentation, per-arm tokenization, skill-conditioned parameters, and cross-arm attention comparable in a way prior systems with different backbones and datasets did not allow. The reported interaction between skill-conditioned parameters and restricted attention is a concrete design signal rather than a single end-to-end result.
Real-world applications.
- Manufacturing and assembly cells where two arms must sometimes act in sequence and sometimes in parallel depending on part arrival and fixture state.
- Household and kitchen robotics where a robot must reorder or synchronize bimanual steps such as holding a container while placing an item into it.
- Warehouse and logistics picking, packing, and sorting, where object and target arguments change constantly while the underlying pick, place, push, and move skills stay the same.
- Laboratory or cabinet-based automation where one arm opens an articulated compartment while the other inserts or retrieves an object, under varying timing and ordering.
Industry relevance. The results are relevant to anyone deploying a shared dual-arm policy rather than two independent single-arm controllers: the independent-control baseline reaches only 5.53% in simulation and 10.00% on the five unseen SO101 conditions, while the single combined architecture reaches 21.53% and 39.00%. The finding that modest architectural changes to an existing pretrained backbone can yield large generalization gains — while costing some in-domain success — directly informs build-versus-buy decisions for bimanual manipulation stacks, and the acknowledgement that pure synchronization remains at 8.67% in simulation identifies where deployment risk concentrates.
Future Directions
- Add a high-level planner. The current study deliberately supplies skill plans via a rule-based scheduler. The authors propose a planner that decomposes novel tasks into familiar skills, assigns them to arms, determines order and synchronization, and revises plans from execution feedback, enabling evaluation of the full planning–execution loop.
- Improve synchronization specifically. Pure Sync remains the weakest group for AE-VLA at 8.67% in simulation, and coordination failures such as arm-wise collisions are the diagnosed cause of Dual π0.5 failures. Closing this gap is an explicit open problem.
- Recover in-domain performance. AWA reduces in-domain CCSR (Token Group falls from 45.00% to 25.00%; AE-VLA reaches 27.17%, below all three baselines, and 60.00% on SO101 versus Dual's 70.00%). Reconciling generalization gains with in-domain retention is unresolved.
- Disentangle the two design choices internally. The authors note that SkillLoRA adds both weights and router supervision, and AWA changes both cross-arm and within-chunk attention, so the study separates the two choices but does not isolate each change inside them.
Target Audience
Robotics and embodied-AI researchers working on bimanual manipulation, VLA policy design, and compositional generalization; benchmark designers who need a controlled protocol for comparing architectural variants under a fixed backbone and dataset; and practitioners evaluating whether a shared dual-arm policy or independent per-arm controllers better suit deployments that will encounter coordination patterns not present in training data. Readers seeking strong prior familiarity with flow-matching action experts, LoRA adaptation, and attention-mask design will get the most from the implementation sections.
Authors’ abstract
Generalization in multi-arm collaboration can be studied as composing familiar atomic skills in new ways across arms. However, existing evaluations offer limited insight into which training and architectural choices support this ability under different coordination requirements. We introduce \textbf{ACG-Bench}, a benchmark for \emph{Arm-wise Compositional Generalization} that provides a common testbed for studying skill recomposition in dual-arm policies. It contains 23 task--condition pairs across 8 task families, with 6 in-domain conditions and 17 unseen compositions covering reordering, synchronization, their combination, and cross-task composition. All methods receive the same per-arm atomic prompts, and success requires achieving the task goal while satisfying physical milestones and specified order or timing constraints. Using $π_{0.5}$ as a common vision-language-action backbone, we compare representative data-augmentation and architectural strategies with shared source data and a common evaluation protocol. Our architectural study examines arm-token grouping, skill-specific LoRA adapters (SkillLoRA), and arm-wise attention (AWA), highlighting the complementarity of skill-conditioned parameters and attention structure. Combining these choices yields \textbf{AE-VLA}, which achieves 21.53\% generalization success in simulation, compared with 2.94\% for Single $π_{0.5}$, 3.06\% for MA-VLA, and 5.53\% for two independently controlled $π_{0.5}$ policies. On physical SO101 robots, AE-VLA reaches 39.00\% mean success across five unseen conditions, compared with 10.00\% for the strongest baseline. These findings provide empirical guidance for designing dual-arm policies that generalize beyond fixed training routines.