Research
A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies
A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action Policies Overview Research area: Safe robot control and deployment-time decoding for vision-language-action (VLA) polic

- arXiv
- 2610.05166
- Published
- 2026-10-04
- Authors
- Tu Nguyen, Matthieu Zimmer, Vu Anh Vu, Ziyi Wang, Jannik Hammel Nielsen, Xuebing Zhou, Haitham Bou Ammar
AI summary
A Safe Action Is Not Enough: Feasible-Future Decoding for Vision-Language-Action PoliciesOverview
Research area: Safe robot control and deployment-time decoding for vision-language-action (VLA) policies, combining constrained decoding, control-as-inference, and safety-critical mobile manipulation benchmarks.
Technical level: Advanced. The paper builds an exact next-block marginal from a history-conditioned policy–environment trajectory law, defines a "feasible-future mass," and proves recovery bounds for a selective finite-candidate approximation, though the practical decoder itself is described as training-free and rollout-free.
Scope: One sentence: the authors show that locally safe or likely actions can still be dead ends under a frozen VLA policy, derive an exact feasible-future target for next-block selection, and introduce the training-free reranker VICS-G, which lowers observed safety cost across six Safety-CHORES settings while keeping success and episode length close to policy sampling.
What This Paper Is About
A frozen vision-language-action policy can prefer a move that is locally admissible yet leaves no policy-supported route to safe task completion. The authors call this mismatch the feasibility–likelihood gap: likelihood ranks the next move, but feasibility depends on the futures that move leaves open. The paper's goal is to bring those futures into the decision without retraining the policy or running online rollouts, and to characterize when a small, selectively applied reranker can recover the ideal choice.
Key Contributions
-
An exact next-block target for policy-relative safe completion. The authors derive the next-action marginal of a prior-preserving trajectory law restricted to safe completion (Theorem 1). The resulting feasible-future mass $Z_{\mathrm{feas}}(\mathcal{H}t,a)$ connects the current action to the safe continuations it leaves open; its support ($\chi{\mathrm{feas}}^{\star}$) records whether safe completion remains possible, and its magnitude ($\overline{Z}_{\mathrm{feas}}^{\star}$) measures how much weighted safe-completion mass remains.
-
A theory of selective approximation and its limits. The authors separate four approximation interfaces — intervention, coverage, support, and ranking — and derive a practical support-separation margin $\Delta_t^{\mathrm{sup}}$, a recovery bound for the best retained viable candidate (Theorem 2), and conditions under which finite continuation rollouts preserve the ideal winner (Corollary 1). The deployed-action loss decomposes into a coverage term, a false-rejection term, and a ranking error $\varepsilon$.
-
A practical route to safer execution with frozen policies. VICS (Viability-Corrected Sampling) is a training-free reranker combining a gate, a likelihood trust region, predicted-safety and task-feasibility filters, a history-dependent local score, and fallback to the policy sample. VICS-G is the generic rollout-free configuration; VICS-S adds task-grounded symbolic-support rules.
-
Empirical validation under realistic operating conditions and imperfect execution. Six Safety-CHORES task/checkpoint comparisons plus matched ablations (zero-weight $\eta_v=0$ sweeps, feedback masking, a "Switch after failure" baseline), and a simulator-assisted rollout extension called VICS-R.
Main Findings
-
Consistent cost reduction across all six settings. VICS-G lowers observed mean safety cost by 1.9%–57.5% across the six Safety-CHORES settings and records fewer violations in every setting, spanning both task-optimized and safety-aligned checkpoints. Success stays within 2.5 percentage points of policy sampling, and mean episode length remains within 0.82 steps.
-
Task progress separates decoders from early-stopping solutions. The RCD baseline attains lower raw cost but its success falls by 3.5–20.3 percentage points and its episodes shorten by 4.4–37.5 steps relative to policy sampling. The paper's motivating ObjectNav episode ("navigate to an apple") shows policy sampling failing at cost 23, RCD ending early at cost 3, and VICS-G reaching the goal at cost 11.
-
VICS-S yields complementary task-specific gains. VICS-S reaches its strongest joint gain on base Fetch: success rises by 3.5 points, cost falls by 46.0%, and violations fall by 41.5%. It matches or improves policy success in four settings and lowers cost in five by 10.3%–47.3%. On safety-aligned Fetch its highest success (58.7%) comes with higher cost and longer episodes (cost reduction of −38.3%).
-
The continuation term $z_G$ adds setting-dependent value. At $\eta_v = 0.10$, matched sweeps record fewer violations in five of six settings and lower mean cost in four than at zero weight; success is unchanged on both PickUp checkpoints and Fetch-safe. ObjectNav-base is the adverse case: cost and violations rise, and success falls by 3.00 points. Fetch sweeps use all 20 actions, so they do not decompose the main top-12 result.
-
Execution feedback preserves near-policy success under perturbed actions. On safety-aligned Fetch with $\rho=0.1$, the feedback-aware VICS-G reaches 55.67% success with 15.9% lower observed cost and 16.6% fewer violations than policy sampling at virtually identical mean episode length.
-
The feedback safety benefit remains uncertain. Against the same decoder with feedback masked, feedback yields 9.7% lower observed cost and a +0.58-point success difference (95% interval [−1.16, 2.33]), meeting the preset 5-point tolerance. Because the cost interval spans zero, confirmatory testing stops before the policy comparisons, and clean-condition cost estimates favor masking.
-
A theoretical separation between local safety and viability is exhibited exactly. Figure 2 traces four candidates — an unsafe shortcut $a_U$, a locally admissible dead end $a_D$, and feasible actions $a_L$, $a_H$ with lower and higher continuation magnitude. Each scoring factor changes the winner: local admissibility removes $a_U$, but only future support removes $a_D$, and feasible-future magnitude then favors $a_H$ over $a_L$.
Methodology in Plain English
The authors start from a law over complete trajectories rather than single actions. They take the joint policy–environment process, weight it by reward and cost, restrict it to trajectories that complete the instruction while satisfying a hard safety budget, and then marginalize out everything after the current block of actions. The algebra leaves a clean product: the policy's own likelihood for the action block, a local reward-minus-cost term, and a factor $Z_{\mathrm{feas}}$ that summarizes all the safe futures the action leaves reachable under the frozen continuation process.
Because that factor cannot be computed exactly online, they build a practical decoder around a small candidate set (proposal budget K = 12, with the policy sample retained, so the admitted set has at most K + 1 members). A gate decides whether the sampled action is worth revisiting — it fires when the sample repeats a failed or collision-linked action, fails terminal admission, or has an explicit hard prohibition. If it fires, admitted candidates are scored with policy log-likelihood plus a local robustness margin plus a bounded continuation-correction term $z_G$ built from cheap monitor and history features. An empty surviving set or a final authorization veto falls back to the policy sample. VICS-S instead uses symbolic eligibility rules to decide which candidates enter the comparison at all, then ranks the survivors with the same generic score.
For evaluation, they run the PickUp, ObjectNav, and Fetch task families of Safety-CHORES with task-optimized and safety-aligned SafeVLA checkpoints, using 160 PickUp, 200 ObjectNav, and 172 Fetch episodes per method, holding policy weights fixed. The single shared VICS-G configuration was selected using roughly 10% of episodes per task family, with reported results over the full sets. A separate study perturbs action magnitudes (IID or episode-persistent) on safety-aligned Fetch and either supplies or masks the preceding commanded-versus-realized transition record. A final, separately baselined study in a modified simulator explores VICS-R, a modest rollout augmentation guided by safe-completion estimates under a limited simulation budget.
Why This Matters
The paper reframes robot safety at deployment time: not merely "is this action safe?" but "is this action viable?" The authors show that a decoder can reduce measured harm while keeping task completion intact, and they give an exact policy-relative target plus explicit error decomposition so that approximate decoders can be audited rather than merely benchmarked.
Real-world applications:
- Mobile manipulation in warehouses or homes, where a robot must reach and grasp a requested object without stalling in a dead end.
- Navigation assistants that currently stop to avoid cost, at the price of never arriving — the behavior the paper highlights in Figure 1.
- Safety-aligned deployed policies where retraining is expensive or impossible, since the method is training-free and keeps policy weights frozen.
- Systems that must report safety and task success jointly, since the authors measure cost, success, and episode length together.
Industry relevance: The work comes from Huawei Heisenberg Research Center, Huawei Noah's Ark Lab, TU Berlin, and UCL Center for AI. Its appeal to industry is operational: a reranker that needs no policy retraining and no online rollouts can be layered on top of existing VLA deployments, and its guarantees are stated relative to the frozen policy's own continuation process rather than an unattainable global optimum.
Future Directions
- A learned safe-completion critic. The authors explicitly name this as another possible approximation to the feasible-future term and leave it to future work.
- Reliable, affordable rollouts. Feasible-future estimates scale poorly to many actions and decision points even in simulation, and their reliability is tied to the fidelity of the dynamics and contact models. The reported VICS-R study is described as a modest rollout augmentation under a limited simulation budget; the content supplied here is truncated inside Section 4.4, so VICS-R's numbers are not reported here.
- Resolving the feedback ambiguity. Under the tested perturbations, execution feedback met the preset success tolerance, but its incremental safety benefit remained uncertain because the cost interval spans zero, and clean-condition estimates favored masking. Determining when realized-transition evidence genuinely helps is an open question.
- The gate's blind spot. A gate that accepts a dead end never reaches the reranker, and the authors note that triggering reranking alone gives no bound on the executed action. Whether dead ends can be detected before intervention, rather than after, remains open.
Target Audience
Researchers and engineers working on safe robot learning, VLA policy deployment, and constrained or guided decoding will benefit most, along with readers interested in control-as-inference derivations and error decompositions for approximate decision rules. The paper assumes comfort with trajectory-level probabilistic modeling, KL-regularized control, and MAP-style decoding; practitioners chiefly interested in applying a drop-in safety layer may find the empirical sections (Sections 4.2–4.3) more immediately useful than the theory.
Authors’ abstract
A safe action is not necessarily a viable one. A frozen vision-language-action (VLA) policy can favor a locally admissible move that leaves no policy-supported route to safe task completion. We call this the feasibility-likelihood gap: likelihood ranks the next move, while feasibility depends on the futures it leaves open. To bring those futures into the decision, we derive the exact next-block marginal of the history-conditioned policy-environment trajectory law restricted to safe task completion. The derivation reveals a candidate-dependent feasible-future mass: its support records whether safe completion remains possible under the frozen continuation process, while its magnitude measures how much weighted safe-completion mass remains. Since exact evaluation is impractical online, we develop a selective finite-candidate approximation and establish conditions for recovering the best retained viable candidate. Our alarm-triggered, training-free reranker VICS-G lowers mean cumulative safety cost by 1.9%-57.5% across six Safety-CHORES settings while remaining within 2.5 percentage points of policy sampling in success and 0.82 steps in mean episode length. Our approach offers a promising and practical path toward safer task completion, grounded in an exact policy-relative target yet requiring neither policy retraining nor online rollouts.