MileTone_2 / models /sft /README.md
SciCode's picture
Add selected SFT variants S1-S4
3eea076 verified
|
Raw History Blame Contribute Delete
1.99 kB
# Selected SFT Checkpoints
This catalogue records the four SFT variants explicitly selected for the Milestone 2
model-only increment. Reported pass rates are historical Runnable alignment results,
not results on the later repository-level Core/Hold benchmark.
| ID | Delivered directory | Initial model | Training data | Final checkpoint | Reported pass@1 | Reported pass@5 |
|---|---|---|---:|---:|---:|---:|
| S1 | `sft_f3_clean_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | F3 clean, 473,465 records | step 14,550, epoch 1 | 9.1% | 13.0% |
| S2 | `sft_f3_hq_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | F3 high-quality v4, 134,775 records | step 4,212, epoch 1 | 8.3% | 13.0% |
| S3 | `sft_f3f4_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | F3 clean + F4 clean, 491,004 records | step 15,033, epoch 1 | 5.7% | 11.5% |
| S4 | `sft_f4_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | F4 clean, 17,539 records | step 484, epoch 1 | 0.9% | 2.5% |
## Authoritative local sources
| ID | Source checkpoint |
|---|---|
| S1 | `/raid/data/weifeng/Pre_train/output_user/qwen25_coder_0p5b_instruct_F3_clean_1epoch_1gpu/v1-20260505-184601/checkpoint-14550` |
| S2 | `/raid/data/weifeng/Pre_train/output_user/qwen25_coder_0p5b_instruct_F3_v4_1epoch_1gpu/v0-20260519-205734/checkpoint-4212` |
| S3 | `/raid/data/weifeng/Pre_train/output_user/qwen25_coder_0p5b_instruct_F3F4_clean_1epoch_1gpu/v0-20260506-103602/checkpoint-15033` |
| S4 | `/raid/data/weifeng/Pre_train/output_user/qwen25_coder_0p5b_instruct_F4_clean_1epoch_1gpu/v3-20260505-141059/checkpoint-484` |
S3 contains 491,004 examples: 473,465 F3 records plus 17,539 F4 records. This exact
count supersedes the approximate 150K value in an earlier historical report.
Each directory contains a complete Hugging Face inference checkpoint: configuration,
tokeniser, chat template, generation configuration, safetensors weights, effective
arguments, and trainer state. Optimiser and scheduler states are intentionally excluded.