Skip to content
AI.info

Research

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing

Overview Research area: Robot learning, specifically vision-language-action (VLA) policies, deployment-time adaptation, and simulation-based robustness benchmarking for manipulation. Technical level:

Taming VLAs under Robot Execution Errors: Self-Compensation and Stress Testing
arXiv
2609.37334
Published
2026-09-29
Authors
Sohyun Lee, Yoonjae Baek, Jaesang Won, Jinnyeong Kim, Kang Hyunwoo, Seung-Hwan Baek, Ivan Laptev, Suha Kwak

AI summary

Overview

  • Research area: Robot learning, specifically vision-language-action (VLA) policies, deployment-time adaptation, and simulation-based robustness benchmarking for manipulation.
  • Technical level: Advanced. The paper assumes familiarity with VLA architectures (flow-matching action experts, LoRA), operational space control, and classical joint-level dynamics models (Stribeck friction, backlash deadband, elastic-joint compliance, gravity-compensation error).
  • Scope: The paper proposes a deployment-time method that lets a VLA pre-compensate for execution errors using command-execution residuals, and introduces RoboStress, a simulation benchmark that injects joint-level error models to stress-test VLA robustness across seven deployment scenarios.

What This Paper Is About

VLA policies typically assume that the motion a robot actually executes matches the action the policy commanded, but real robots deviate because of wear, heavy payloads, temperature changes, gear clearance, and joint compliance. Existing robustness studies mostly perturb the policy's inputs (viewpoints, lighting, paraphrased instructions) or inject synthetic noise directly into end-effector commands, which does not reflect how execution errors physically arise. The paper's goal is to (1) make a deployed VLA adapt online so its commands pre-compensate for the errors of the specific robot it runs on, and (2) provide a controlled simulation benchmark to test that robustness across realistic, state- and history-dependent execution conditions.

Key Contributions

  1. Self-compensating VLA: a deployment-time adaptation method that updates a pretrained VLA online using the residual between the commanded delta pose and the proprioceptively measured executed motion, without task rewards, labels, or knowledge of the error sources.
  2. A compensated pseudo-target and objective: the observed residual is subtracted from the issued command to form a target action chunk, and only rank-4 LoRA adapters on the action expert are updated with a flow-matching loss plus an anchor term, leaving the pretrained weights and the low-level controller unchanged.
  3. RoboStress: a controlled simulation benchmark that combines established joint-level models of Stribeck friction, gravity-compensation error, backlash, and dynamic compliance, injects each at the corresponding stage of the control pipeline, and organizes them into seven deployment scenarios.
  4. Validation in simulation and on hardware: evaluation across all seven RoboStress scenarios and on two physical AgileX Piper arms with different usage histories, including generalization to objects unseen in the fine-tuning demonstrations.

Main Findings

  • RoboStress success rates: self-compensating VLA outperforms the base policy, domain randomization (DR), and RobustVLA in all seven deployment scenarios with both backbones, raising the average (excluding Clean) from 49.1% to 56.6% with π0.5 and from 41.7% to 44.2% with π0.
  • Training-time robustness methods are inconsistent: DR and RobustVLA fall below the base policy on average with π0 (41.7% base versus 39.8% for DR and 38.8% for RobustVLA), while gaining modestly with π0.5 (49.1% base, 51.0% DR, 49.8% RobustVLA, 56.6% Ours).
  • Clean-setting cost: on the Clean scenario the method scores 95.8% with π0.5 (base 97.0%) and 92.6% with π0 (base 91.1%), so the gain is concentrated in the perturbed scenarios rather than the clean one.
  • Real robots: on a new Piper arm (Robot A) and an arm used for one year (Robot B), self-compensating VLA beats the base policies and RobustVLA on every task with both π0 and π0.5; averages with π0.5 are 76% (Robot A) and 74% (Robot B) versus base averages of 43% on each arm.
  • Both arms show residual mismatch: even the new arm exhibits command-execution mismatch, with a mean normalized residual of 32.9%, compared with 35.0% on the older arm.
  • Real-world gains exceed 30 percentage points: the abstract reports that the method raises the average task success rate by more than 30 percentage points on each arm for both base policies.
  • Per-component analysis: the method improves the average from 50.0% to 51.5% across the individual noise settings (Stribeck 36.8% to 37.8%, Backlash 68.0% to 72.1%, Compliance 74.1% to 74.2%, Gravity 21.1% to 22.0%); the smallest gain is under compliance, where the base policy already achieves its highest success rate.
  • Both ingredients are necessary: ablating the residual (setting η_obs = 0) drops performance to 40.8%, and direct residual correction without policy adaptation reaches only 23.6%, both below the full method (44.2%) and, in the direct-correction case, below the base policy (41.7%).
  • Complementary to external correction: on Task 1 of Robot A with π0.5, a disturbance observer (DOB) alone lifts success from 48% to 56%, adding the adaptation reaches 80%, compared with 84% for the adaptation alone.
  • Simulation fidelity check: when a pick-and-place demonstration is replayed on a Piper arm carrying a bowl with 1 kg of weights, the adjusted RoboStress Heavy Payload model captures the post-grasp rise in deviation and yields a lower dynamic time warping (DTW) distance to the real mean profile than Gaussian action noise.
  • Generalization to unseen objects: on Task 1 of Robot A with five unseen object variants (including wooden, plum-colored, and light-blue bowls, 25 episodes total), the method reaches 64% with π0.5 versus 16% for the base policy and 36% for RobustVLA.

Methodology in Plain English

The approach treats the mismatch between what the robot is told to do and what it actually does as a free training signal. At each executed step, the system compares the commanded change in end-effector pose (or joint position, on the Piper arms) with the change measured through proprioception, producing a residual. It then forms a target command by subtracting that residual from the command it issued, meaning "next time, ask for a little more or less than before so the robot lands where intended." Because the residual is only known after the motion happens, this target cannot correct the current action chunk — it is used to update the policy for subsequent chunks. Updates are restricted to small rank-4 LoRA adapters on the action expert of a frozen base policy, trained with a flow-matching loss over a buffer of accumulated target chunks plus a penalty keeping the adapters close to their initialization, so the low-level controller stays untouched.

For stress testing, the authors built RoboStress in simulation. It takes four well-established joint-level error models — Stribeck friction, gravity-compensation error, backlash (a deadband on direction reversal), and compliance (a spring-mass-damper deflection under load) — and injects each where it physically belongs: friction and gravity errors as torques added to the controller output, backlash and compliance as offsets to joint positions feeding physics integration. Because the joints are coupled through the mass matrix, a perturbation at one joint propagates to the others, so the residual a policy sees is not the injected noise itself but its effect at the current configuration. Severity factors scale the model parameters, and seven scenarios combine the components to mimic heavy payloads, thermal drift, and wear on different joint sets. The method is then compared against the two base policies, domain randomization on the action space, and adversarial RobustVLA training.

Why This Matters

  • Impact on research: the paper reframes robustness for VLA policies away from input corruptions and synthetic end-effector noise toward execution errors that arise from robot mechanics and operating conditions, and it offers both a method and a reproducible benchmark for that setting.
  • Manufacturing and assembly: arms with worn gears, degraded lubrication, or varying payloads could keep working without manual recalibration, since compensation happens at the policy level without changing the controller or requiring an explicit dynamics model.
  • Warehouse and logistics robots: fleets of nominally identical arms differ in usage history, and per-robot online adaptation addresses that heterogeneity without collecting per-robot labels or rewards.
  • Field and service robotics: temperature changes, such as thermal drift over an episode, are handled by an approach that updates continuously rather than assuming stationary conditions.
  • Destructive-testing reduction: RoboStress allows stress testing of extreme execution conditions in simulation, reducing the need to expose physical hardware to potentially damaging settings.

Industry relevance: the method modifies only adapters on a pretrained policy and leaves the existing controller unchanged, which makes it a comparatively low-risk addition to deployed VLA stacks; the paper also notes that reported gains are not safety guarantees and that human-environment deployment requires additional safety evaluation.

Future Directions

  1. Reducing reliance on proprioception: the authors explicitly list incorporating motion estimates from external sensors as future work, which would help when joint feedback is noisy or unavailable.
  2. Generalizing beyond the tested hardware: the real-world evaluation covers two AgileX Piper arms and four tasks, leaving open how the method scales to other embodiments, degrees of freedom, and gripper types.
  3. Combining with external controllers: results with a disturbance observer suggest complementarity rather than redundancy, raising the question of how best to blend policy-level adaptation with controller-level correction and adaptive control.
  4. Broadening the benchmark: RoboStress covers four joint-level noise components across seven scenarios; extending coverage to additional error sources and closing the remaining gap between simulated and real deviations, as probed by the DTW comparison, are natural next steps.

Target Audience

Researchers and engineers working on robot manipulation, VLA policies, and sim-to-real transfer will get the most from this paper, since it assumes comfort with flow-matching policies, LoRA fine-tuning, and joint-level dynamics models. It is also relevant to practitioners deploying learned policies on physical arms who need robustness to wear, payload, and thermal variation, and to benchmark designers interested in simulating execution errors rather than injecting synthetic command noise. Readers looking for a beginner-level introduction to VLA robustness will find the equations in Sections 3 and 4 demanding.

Authors’ abstract

Vision-language-action (VLA) policies often fail when a robot's executed motion deviates from their commanded action. Such execution errors arise from the robot's mechanics and operating conditions, such as wear and payload changes. We propose self-compensating VLA, a deployment-time adaptation method that enables a VLA policy to pre-compensate for the robot's execution errors when generating commands. Without task rewards or labels, it updates the policy online using the residual between the action commanded by a VLA and the motion executed by the robot. To stress-test VLA robustness across execution conditions that are impractical to cover with physical robots alone, we introduce RoboStress, a controlled simulation benchmark. It combines established joint-level models of friction, backlash, compliance, and gravity-compensation error into seven deployment scenarios whose execution errors depend on the robot's state and motion history. On RoboStress, self-compensating VLA achieves higher average task success than both the base policies and methods that build in robustness during training. On two physical robot arms with different usage histories, it raises the average task success rate by more than 30 percentage points on each arm, and the gains extend to objects not seen in the task demonstrations.

Read the original paper