Yuki131's picture
Create README.md
361f9eb verified
|
Raw History Blame Contribute Delete
2.68 kB

R2 Training Checkpoints (Stages 1–3)

We release the checkpoints from the three-stage training pipeline described in the third version of our paper. Stage 1 uses supervised fine-tuning; Stage 2 produces two checkpoints through soft-label distillation; and Stage 3 combines them through model soup to produce the final R2 models.