VIP Qwen3-4B math โ classification value, AP, active sampling
Public training checkpoints for the VIP math run tracked at Weights & Biases.
Each completed 100-step checkpoint is published as a branch named step-N after tokenizer/config sanitization. The run uses 32 prompts per update, group size 8, answer-prefix value conditioning, active sampling, zero-variance filtering, 8 async steps, classification value loss, and one critic epoch.
Model tree for hamishivi/vip-g8-p32-cls-ap-as-zvf-qwen3-4b-math
Base model
Qwen/Qwen3-4B-Base