Skip to content
AI.info

Research

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

Overview Research area: Robotics — Vision-Language-Action (VLA) policies, inference-time (test-time) scaling, and learned verifiers for action selection. Technical level: Intermediate. The paper assum

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
arXiv
2610.04933
Published
2026-10-04
Authors
Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe, Namiko Saito, Katsushi Ikeuchi, Sharon Li, Yasuyuki Matsushita

AI summary

Overview

  • Research area: Robotics — Vision-Language-Action (VLA) policies, inference-time (test-time) scaling, and learned verifiers for action selection.
  • Technical level: Intermediate. The paper assumes familiarity with VLA models, Best-of-N sampling, binary discriminators, and representation-space analysis, though the core idea is explained with accessible intuition.
  • Scope: The paper proposes DiVeR, a decision-criticality-weighted verifier that concentrates verifier training on sparse, consequential states, and evaluates it on LIBERO-Long, RoboCasa, and a real Franka Research 3 robot with the π0 and π0.5 policies.

What This Paper Is About

VLA robot policies can be improved at inference time without new data by sampling multiple candidate action chunks and using a learned verifier to pick the best one (Best-of-N). Existing classification-based verifiers learn from whole-trajectory success or failure labels but weight every visited state equally, even though most states are routine and offer little signal for choosing between actions. The goal of this work is to reweight verifier training so that learning concentrates on the sparse decision-critical states where the choice between candidate actions actually determines whether the task succeeds.

Key Contributions

  1. Diagnosis of uniform verifier training. The authors identify that routine states dominate trajectories and provide limited discrimination signal, while a sparse set of decision-critical states (such as pre-grasp or pre-placement alignment) is far more consequential for action selection. Figure 1(b) shows this sparsity across 500 LIBERO-Long trajectories from a frozen π0 policy using the criticality measure u_t.
  2. The DiVeR method. A decision-criticality-weighted verifier that uses the dispersion (mean per-coordinate variance) of action-expert representations sampled from the frozen VLA as a practical proxy for decision criticality, requiring no step-level annotations and no additional environment interaction.
  3. Broad empirical validation. Consistent gains on LIBERO-Long, RoboCasa, and four real-world Franka Research 3 tasks with π0 and π0.5, using a single verifier shared across all tasks within each benchmark.
  4. Efficiency and generalization evidence. DiVeR outperforms a VLM-based verifier (RoboMeter) while being over 700 times faster at verification, and it transfers across held-out task categories.

Main Findings

  • Headline gains: Across policies and environments, DiVeR improves task success by up to +13.6% over single-sample inference and +3.4% over the strongest baseline. It surpasses the VLM-based verifier by +4.6% success rate while reducing verifier inference latency by over 700 times.
  • LIBERO-Long (Table 1): With π0 and N = 32, DiVeR improves success rate by +5.8% over the N = 1 baseline (83.5) and by +2.8% over the best-performing baseline, reaching 89.3 ± 0.3. With π0.5, DiVeR reaches 94.8 ± 0.3 at N = 4, 96.0 ± 0.5 at N = 8, 97.0 ± 0.5 at N = 16, and 97.5 ± 0.5 at N = 32.
  • RoboCasa (Table 2): At N = 16, DiVeR improves average success rate by +5.9% for π0 (38.1 vs. baseline 32.2) and +11.6% for π0.5 (54.7 vs. baseline 43.1), improving across articulated, control, and object-centric categories rather than favoring one task type. Navigation success remains 0.0 for the baseline and most methods, and DiVeR reports 15.0 for π0.5.
  • Real robot (Table 3): On a Franka Research 3 with N = 16, DiVeR reaches 71.9% average success over 24 episodes per task, versus 65.6% for a uniformly trained verifier and 58.3% for the N = 1 base policy — gains of +6.3 and +13.6 percentage points. Per-task results for DiVeR are 75.0 (PnP block), 70.8 (PnP cup), 66.7 (stack block), and 75.0 (close drawer).
  • Component analysis (Table 4): Random selection gives only marginal gains (33.0 for π0, 44.1 for π0.5) over the N = 1 baseline (32.2, 43.1), showing samples help only with a good selection rule. Uniform training on action-expert representations improves by +2.7 and +7.5 points over baseline, and further improves over the original SVM by +1.0 and +3.7 points. Decision-criticality weighting adds +3.2 and +4.1 points on top of uniform training, yielding 38.1 and 54.7.
  • Input feature ablation (Table 5): Action-expert representations alone (38.1 for π0, 54.7 for π0.5) outperform raw actions and match or beat image-plus-representation inputs (37.4, 54.6), so image features provide no additional benefit.
  • VLM comparison (Table 6): RoboMeter achieves 50.1% success with 0.743 s latency, versus DiVeR at 54.7% with 0.001 s latency — a +4.6 percentage point improvement and over 700 times lower verification latency.
  • Cross-task generalization (Figure 4): Trained on one of three functional RoboCasa categories (articulated, control, object-centric) and evaluated on held-out tasks, DiVeR transfers effectively across categories. Navigation tasks were excluded due to qualitatively different behaviors.
  • Architecture choice (Figure 6): Under a matched budget of approximately 66K trainable parameters each, an MLP verifier outperforms LSTM and single-layer transformer verifiers on both VLA policies, suggesting the frozen action-expert representation is already sufficient for discrimination.
  • Qualitative behavior (Figure 5): From identical decision-critical states, a uniformly trained verifier selects actions leading to failure (gripper slightly above the drawer front; red block placed beside rather than on the blue block), while DiVeR selects actions that complete drawer closing and block stacking.

Methodology in Plain English

The base VLA policy is kept frozen and treated as a stochastic generator: at each decision step it can produce several candidate action chunks. A small learned verifier then scores candidates and the highest-scoring one is executed. To train that verifier, the authors collect offline trajectories with success or failure labels and propagate the trajectory outcome to every timestep, as in prior work.

The distinctive step is how states are weighted. At each visited state, the frozen policy samples K = 32 candidate action chunks, and the corresponding internal hidden representations are extracted from the policy's action expert. The mean per-coordinate variance of those representations, u_t, serves as the "decision criticality" score: low variance means candidates look alike (routine states), high variance means the policy's plausible actions genuinely diverge (decision-critical states). These scores are standardized within each trajectory, converted into positive weights via exponentiated weighting w_t = exp(β·ũ_t), and used to weight a binary cross-entropy objective that separates successful from failed visitation. The weights depend only on candidate disagreement, not on the outcome label, so no step-level annotation is required.

At inference, the verifier scores N sampled candidates at each step and the best is executed. The verifier is a two-layer MLP with hidden dimension 64 and dropout 0.3 over final-layer action-token representations, trained for 30 epochs with AdamW (learning rate 10⁻³, weight decay 10⁻⁴, batch size 512, class-balanced sampling), with β = 0.5 and weights clipped to [c⁻¹, c] with c = 2.5. All training and inference ran on a single NVIDIA RTX A6000 GPU with 48 GB. Hyperparameters were tuned on RoboCasa with π0.5 and then fixed across benchmarks and policies. Baselines include the verifier-free KDPE and MG-Select, the learned verifiers TACO and SVM, and the VLM-based RoboMeter.

Why This Matters

  • Impact on research: The paper shifts attention from "generate more candidates" to "learn better where to trust the verifier." It connects VLA test-time scaling to the observation in language modeling that only a sparse subset of high-entropy "forking" tokens steers downstream outcomes, and it provides a proxy for state importance that needs neither step-level labels nor extra environment interaction.
  • Real-world applications:
    • Industrial and warehouse manipulation, where verifier-guided selection can improve pick-and-place reliability without new demonstrations.
    • Household service robots performing drawer closing, cup and plate placement, and block stacking, the four tasks evaluated on the Franka Research 3.
    • Deployment settings with tight latency budgets, where a 0.001 s verifier is practical and a 0.743 s VLM judge is not.
    • Multi-task robot fleets, since a single verifier is shared across all tasks in each benchmark and transfers to held-out task categories.
  • Industry relevance: The method requires only that a policy sample multiple action candidates and expose an internal action representation, so it can be layered on existing deployed policies as an inference-time wrapper; the negligible verifier overhead and the ability to reuse VLM computation through KV caching make it attractive for cost-sensitive real-time control.

Future Directions

  • Extending decision-criticality weighting to navigation tasks, which were excluded from the zero-shot generalization study due to their qualitatively different behaviors.
  • Applying the method to policies beyond flow-matching architectures; the authors evaluate the autoregressive OpenVLA in Appendix C.2, indicating this is an open generalization question.
  • Investigating the weighting design itself — the temperature β, clipping threshold c, and the choice to standardize criticality within each trajectory — and whether other criticality proxies correlate as well with rollout-based oracle criticality (a correlation checked in Appendix C.1).
  • Combining this parallel Best-of-N formulation with sequential test-time scaling, such as extended rollouts or iterative refinement, which the paper explicitly sets outside its definition of test-time scaling.
  • Reducing the cost of the criticality estimate, which requires K = 32 sampled candidates at every training state, and examining how sensitive results are to that budget.

Target Audience

Robotics and embodied-AI researchers working on VLA policies, test-time scaling, and reward or verifier models; practitioners at robotics companies who want inference-time improvements without collecting new demonstrations or retraining base policies; and machine learning researchers interested in how non-uniform supervision value across a trajectory mirrors findings on forking tokens in language generation. Readers without a background in imitation learning or sampling-based selection will find the results accessible but should expect the method section to require prior familiarity with VLA architectures and binary discriminator objectives.

Authors’ abstract

Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.

Read the original paper