DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
- Seongheon Park ,
- Heecheol Kim ,
- Shulin Tian ,
- Lilika Makabe ,
- Namiko Saito ,
- Katsushi Ikeuchi ,
- Sharon Li ,
- Yasuyuki Matsushita
arXiv
Preprint

Abstract
Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.
Method Overview

We propose DiVeR, a decision-criticality-weighted verifier for VLA test-time scaling. DiVeR learns from trajectory-level success and failure outcomes while keeping the base VLA policy frozen.
- Identify where action selection matters. We estimate decision criticality from the variance of sampled action representations, without requiring additional environment interaction or step-level supervision.
- Focus verifier learning on critical states. We use these criticality estimates to assign greater training weight to states where candidate actions diverge and accurate selection is more consequential.
- Select actions efficiently at test time. The trained verifier scores multiple action candidates and selects the highest-scoring action chunk for execution, enabling Best-of-N action selection with negligible verifier overhead.
Quantitative Results
We evaluate DiVeR with π₀ and π₀.₅ on two simulated manipulation benchmarks, LIBERO-Long and RoboCasa, and with π₀.₅ on a real-world Franka Research 3 robot. We also evaluate DiVeR with the autoregressive policy OpenVLA on LIBERO-Long.
Across these settings, DiVeR consistently improves task success through verifier-guided Best-of-N action selection. Additional analyses examine scaling with the number of action candidates, component ablations, generalization to unseen tasks, and inference latency, demonstrating the effectiveness of decision-criticality weighting with negligible overhead from verifier scoring.



Qualitative Results
We qualitatively compare the base VLA policy with DiVeR on a Franka Research 3 robot. Starting from the same initial state, the base policy may produce imprecise grasping or placement actions that lead to failure, whereas DiVeR selects more appropriate actions and successfully completes the task.
Close the drawer


Stack the red block on the blue block

