DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling
Abstract
Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.
Community
DiVeR improves VLA test-time scaling by identifying sparse decision-critical states from action-representation dispersion and focusing verifier learning where action selection matters most.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Looking Back to Move Forward: Temporal Verification for Generative Robot Policies (2026)
- WISE: World-model-guided Imagination Scheduling for Efficient Post-training of Vision-Language-Action Models (2026)
- Prioritized Rollouts for Efficient World Model-based Vision-Language-Action Policy Optimization (2026)
- eRLT: Efficient VLA Reinforcement Learning via Action-Relevant Token Routing (2026)
- RoboIRS: Inference-Time Internal Representation Steering for Generalist Robot Policies (2026)
- TaskAnchor: Grounding Task State in Reactive VLAs for Long-Horizon Manipulation (2026)
- EvoMem-VLA: State-Evolution Memory for Long-Horizon Robot Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.04933 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper