ViGoRL V* multi-turn RL โ€” 3B ablation checkpoints (seed 42)

Reproduction of ViGoRL (arXiv:2505.23678) multi-turn GRPO on V* (visual search) with five rollout-time ablations, all from the same SFT init and seed. One folder per ablation:

Folder Ablation
vstar_3b_seed42_baseline_step250 original ViGoRL multi-turn (zoom-in crop)
vstar_3b_seed42_no_crop_step250 full uncropped image fed back
vstar_3b_seed42_black_image_step250 all-black canvas instead of the crop
vstar_3b_seed42_no2_noviz_step250 text-only acknowledgement, no image returned
vstar_3b_seed42_no23_multi_round_step250 original image + "use the grounding thoughts above"
vstar_3b_seed42_no123_rethink_step250 text-only rethink, no tool calls, no visual feedback

Training

  • Base: gsarch/ViGoRL-Multiturn-MCTS-SFT-3b-Visual-Search, verl/EasyR1 multi-turn GRPO, vLLM 0.8.3 rollout.
  • Data: vigorl_SA_RL (V*), seed 42, 8 GPUs, rollout batch 128 x n=8, global batch 64, max response 4096, crop 672, offset 182, LR 1e-6 constant after 5% warmup, KL 1e-2, max grad norm 0.2, vision tower frozen.
  • Stopped at step ~250 of the 500-step schedule (250 for all; a few ran to 258-288 before stopping).
  • vLLM prefix caching disabled. The V1 engine enables it by default and it corrupts multi-turn rollouts that append image crops between turns; with it on, tool calling degrades badly at step 0.
  • Rollout-format fixes relative to the upstream branch: newline before </observation>, the V* system prompt made byte-identical to the MCTS-SFT prompt, and the forced-answer string masked out of the loss.
  • Code: williamium3000/demystify-twi, branch vigorl, commit cc1882a.

Training metrics at the stopping point

Format reward ceiling is 1.60 (well-formed rollout + four distinct search coordinates).

Ablation format accuracy (train, letter EM)
baseline 1.60 0.77
no_crop 1.60 0.78
black_image 1.60 0.77
no2_noviz 1.60 0.78
no23_multi_round 1.60 0.74
no123_rethink 1.00 (its own no-tool-call reward) 0.77

All five tool-calling ablations sit at the ceiling, i.e. essentially every rollout issues four valid, distinct search calls.

V*Bench (191 items, step 25 only)

Evaluated with eval_vstar.py (HF Transformers, 4 search turns, temperature 0.5, single sample). These models answer with MCQ letters, because the training reward scores letters, so plain string match against the option text undercounts them; the letter-aware column maps the letter back to its option.

Checkpoint raw string match letter-aware tool calls
no2_noviz step25 0.403 0.644 4/4 on 191/191
no123_rethink step25 0.644 0.675 none (by design)
no_crop step25 0.309 0.639 4/4
baseline step25 0.194 0.586 4/4
black_image step25 0.162 0.586 4/4
no23_multi_round step25 0.215 0.571 4/4
SFT init (reference) 0.696 0.707 4/4 on 82%
official ViGoRL-Multiturn-3b (reference) 0.764 0.780 4/4

At step 25 every ablation is still below the SFT init. Step-250 evaluations are not done yet, so the uploaded checkpoints are not benchmarked; train-time accuracy rose from ~0.45 to ~0.77 after step 25.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for DemystifyingTWI/vigorl-ckpts

Finetuned
(1)
this model

Paper for DemystifyingTWI/vigorl-ckpts