DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

arXiv

Preprint

Comparison of routine and decision-critical states, showing that high decision criticality corresponds to greater action-candidate divergence and occurs sparsely across LIBERO-Long trajectories.

Abstract

Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.

Method Overview

Diagram of the DiVeR framework showing decision-criticality estimation from VLA action candidates, decision-weighted verifier training, and verifier-guided action selection at inference.

We propose DiVeR, a decision-criticality-weighted verifier for VLA test-time scaling. DiVeR learns from trajectory-level success and failure outcomes while keeping the base VLA policy frozen.

  • Identify where action selection matters. We estimate decision criticality from the variance of sampled action representations, without requiring additional environment interaction or step-level supervision.
  • Focus verifier learning on critical states. We use these criticality estimates to assign greater training weight to states where candidate actions diverge and accurate selection is more consequential.
  • Select actions efficiently at test time. The trained verifier scores multiple action candidates and selects the highest-scoring action chunk for execution, enabling Best-of-N action selection with negligible verifier overhead.

Quantitative Results

We evaluate DiVeR with π₀ and π₀.₅ on two simulated manipulation benchmarks, LIBERO-Long and RoboCasa, and with π₀.₅ on a real-world Franka Research 3 robot. We also evaluate DiVeR with the autoregressive policy OpenVLA on LIBERO-Long.

Across these settings, DiVeR consistently improves task success through verifier-guided Best-of-N action selection. Additional analyses examine scaling with the number of action candidates, component ablations, generalization to unseen tasks, and inference latency, demonstrating the effectiveness of decision-criticality weighting with negligible overhead from verifier scoring.

Success rates on LIBERO-Long comparing action-selection methods across policies and numbers of sampled action candidates.
Average success rates (%) on LIBERO-Long with varying numbers of action candidates NNN sampled at test time. Bold and underlined values denote the best and second-best results, respectively.
Category-wise success rates on RoboCasa atomic tasks comparing baseline methods with ours under two policies.
Category-wise average success rates (%) on RoboCasa atomic tasks. Bold and underlined values denote the best and second-best results, respectively.
Real-robot evaluation showing the robot setup, four manipulation tasks, and success rates for Base, Uniform, and Ours.
Real-robot evaluation with π0.5. Left: robot setup. Center: the four evaluation tasks. Right: task success rates (%) over 24 evaluation episodes per task.

Qualitative Results

We qualitatively compare the base VLA policy with DiVeR on a Franka Research 3 robot. Starting from the same initial state, the base policy may produce imprecise grasping or placement actions that lead to failure, whereas DiVeR selects more appropriate actions and successfully completes the task.

Close the drawer

Stack the red block on the blue block