ViGoRL V* multi-turn RL โ 3B ablation checkpoints (seed 42)
Reproduction of ViGoRL (arXiv:2505.23678) multi-turn GRPO on V* (visual search) with five rollout-time ablations, all from the same SFT init and seed. One folder per ablation:
| Folder | Ablation |
|---|---|
vstar_3b_seed42_baseline_step250 |
original ViGoRL multi-turn (zoom-in crop) |
vstar_3b_seed42_no_crop_step250 |
full uncropped image fed back |
vstar_3b_seed42_black_image_step250 |
all-black canvas instead of the crop |
vstar_3b_seed42_no2_noviz_step250 |
text-only acknowledgement, no image returned |
vstar_3b_seed42_no23_multi_round_step250 |
original image + "use the grounding thoughts above" |
vstar_3b_seed42_no123_rethink_step250 |
text-only rethink, no tool calls, no visual feedback |
Training
- Base:
gsarch/ViGoRL-Multiturn-MCTS-SFT-3b-Visual-Search, verl/EasyR1 multi-turn GRPO, vLLM 0.8.3 rollout. - Data:
vigorl_SA_RL(V*), seed 42, 8 GPUs, rollout batch 128 x n=8, global batch 64, max response 4096, crop 672, offset 182, LR 1e-6 constant after 5% warmup, KL 1e-2, max grad norm 0.2, vision tower frozen. - Stopped at step ~250 of the 500-step schedule (250 for all; a few ran to 258-288 before stopping).
- vLLM prefix caching disabled. The V1 engine enables it by default and it corrupts multi-turn rollouts that append image crops between turns; with it on, tool calling degrades badly at step 0.
- Rollout-format fixes relative to the upstream branch: newline before
</observation>, the V* system prompt made byte-identical to the MCTS-SFT prompt, and the forced-answer string masked out of the loss. - Code:
williamium3000/demystify-twi, branchvigorl, commitcc1882a.
Training metrics at the stopping point
Format reward ceiling is 1.60 (well-formed rollout + four distinct search coordinates).
| Ablation | format | accuracy (train, letter EM) |
|---|---|---|
| baseline | 1.60 | 0.77 |
| no_crop | 1.60 | 0.78 |
| black_image | 1.60 | 0.77 |
| no2_noviz | 1.60 | 0.78 |
| no23_multi_round | 1.60 | 0.74 |
| no123_rethink | 1.00 (its own no-tool-call reward) | 0.77 |
All five tool-calling ablations sit at the ceiling, i.e. essentially every rollout issues four valid, distinct search calls.
V*Bench (191 items, step 25 only)
Evaluated with eval_vstar.py (HF Transformers, 4 search turns, temperature 0.5, single sample). These
models answer with MCQ letters, because the training reward scores letters, so plain string match
against the option text undercounts them; the letter-aware column maps the letter back to its option.
| Checkpoint | raw string match | letter-aware | tool calls |
|---|---|---|---|
| no2_noviz step25 | 0.403 | 0.644 | 4/4 on 191/191 |
| no123_rethink step25 | 0.644 | 0.675 | none (by design) |
| no_crop step25 | 0.309 | 0.639 | 4/4 |
| baseline step25 | 0.194 | 0.586 | 4/4 |
| black_image step25 | 0.162 | 0.586 | 4/4 |
| no23_multi_round step25 | 0.215 | 0.571 | 4/4 |
| SFT init (reference) | 0.696 | 0.707 | 4/4 on 82% |
| official ViGoRL-Multiturn-3b (reference) | 0.764 | 0.780 | 4/4 |
At step 25 every ablation is still below the SFT init. Step-250 evaluations are not done yet, so the uploaded checkpoints are not benchmarked; train-time accuracy rose from ~0.45 to ~0.77 after step 25.