YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Self-HPO P-step-3 RL (combo_v1fix), iter 34 (Megatron torch_dist, resumable)
Intermediate slime/Megatron distributed checkpoint (model + optimizer state) of the iteration-3
P-step of Self-HPO (Qwen3.5-35B-A3B, init = P2 sweagent/selfharmo-mstep2-rl-iter49,
harness = H3 rd_c3_combo_v1fix, GRPO, 50 steps planned), so training can resume from iteration 34.
Layout matches the slime --save/--load dir:
iter_0000034/torch_dist shards + metadata.jsonlatest_checkpointed_iteration.txt(= 34)rollout/global_dataset_state_dict_34.pt(data loader position)
Original save dir name: prgrepalign-grpodyn-pe1-r32_n4_max8_s8_g128_grpo_lr1e-6_1nodes_iter3_combov1fix_16k
(under selfharmo-p2-rl-combov1fix/). Download into that dir and point --load at it;
keep RANDOM_SUFFIX=iter3_combov1fix_16k.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support