Warm-started Checkpoints Collection A collection of three models trained on the Nemotron Post Training Dataset for reasoning tasks with IVON ⢠4 items ⢠Updated Aug 25
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published Aug 10 ⢠8
Parameter Exploration for RLVR via Variational Learning Paper ⢠2608.09805 ⢠Published Aug 10 ⢠8
3PO Models Collection 3PO family methods trained on DapoMath-17k using Olmo3-IVON-SFT-7B and Qwen2.5Math-IVON-SFT-7B ⢠10 items ⢠Updated Aug 13