|
Download README.md from SciCode/MileTone_2: direct link, hf CLI and curl.
- Browser
- Download file 6.19 kB
-
https://huggingface.co/SciCode/MileTone_2/resolve/main/README.md
- Command line
-
hf download hf://SciCode/MileTone_2/README.md
-
curl -L -o README.md https://huggingface.co/SciCode/MileTone_2/resolve/main/README.md
6.19 kB
| # Milestone 2 Model and Training Artefact Delivery | |
| This English-language package consolidates nine historical model checkpoints, the | |
| core training and data-preparation code, launch scripts, provenance evidence, datasets | |
| already selected for delivery, and integrity tooling. Four additional SFT checkpoints | |
| (S1--S4) were selected for this model-only increment. | |
| ## Delivered Models | |
| | Model | Path | Initialisation | Training record | | |
| |---|---|---|---| | |
| | CodeGPT multilingual CPT | `models/pretrained/codegpt_multilingual_5epoch/` | `AISE-TUDelft/CodeGPT-Multilingual`, revision `f373547d1234992a64cd08d64cd69abbb3a83579` | 5 epochs, 222,615 steps, 8,753,591,494 training tokens | | |
| | GPT-2 124M CPT | `models/pretrained/gpt2_124m_code_5epoch/` | `openai-community/gpt2`, revision `607a30d783dfa663caf39e06633721c8d4cfcd7e` | 5 epochs, 222,615 steps, 8,753,591,494 training tokens | | |
| | Qwen2.5-Coder-7B CPT | `models/pretrained/qwen25_coder_7b_cpt/` | `Qwen/Qwen2.5-Coder-7B`, revision `0396a76181e127dfc13e5c5ec48a8cee09938b02` | full-parameter CPT, 3 epochs, 18,165/18,165 steps | | |
| | Qwen2.5-Coder-0.5B Stage-1 CPT | `models/pretrained/stage1_cpt/` | `Qwen/Qwen2.5-Coder-0.5B` | continued-pre-training deliverable | | |
| | Qwen2.5-Coder-0.5B-Instruct F3 SFT | `models/sft/sft_f3_refined_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | reported pass@1 9.1%; pass@5 13.0% | | |
| | S1: Qwen2.5-Coder-0.5B-Instruct, F3 clean | `models/sft/sft_f3_clean_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | 473,465 records; step 14,550; reported pass@1 9.1%, pass@5 13.0% | | |
| | S2: Qwen2.5-Coder-0.5B-Instruct, F3 high-quality v4 | `models/sft/sft_f3_hq_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | 134,775 records; step 4,212; reported pass@1 8.3%, pass@5 13.0% | | |
| | S3: Qwen2.5-Coder-0.5B-Instruct, F3+F4 clean | `models/sft/sft_f3f4_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | 491,004 records; step 15,033; reported pass@1 5.7%, pass@5 11.5% | | |
| | S4: Qwen2.5-Coder-0.5B-Instruct, F4 clean | `models/sft/sft_f4_instruct/` | `Qwen/Qwen2.5-Coder-0.5B-Instruct` | 17,539 records; step 484; reported pass@1 0.9%, pass@5 2.5% | | |
| All nine directories are standard Hugging Face checkpoints. The 7B selection is the | |
| completed full-parameter continued-pre-training checkpoint, not the separate historical | |
| 7B LoRA instruction-tuning experiment. | |
| The S1--S4 identifiers, exact source checkpoints, and evaluation results are catalogued | |
| in `models/sft/README.md`. In particular, S3 used 491,004 records (473,465 F3 records plus | |
| 17,539 F4 records); this corrects the approximate 150K count in an earlier historical | |
| report. The checkpoint arguments, launch record, and measured input files support the | |
| 491,004-record count. | |
| Model weights, tokenisers, and available effective-argument records are included. | |
| Optimizer, scheduler, and random-number-generator states are excluded because this is a | |
| model and training-recipe delivery, not a mid-run resume package. | |
| ## Training Code and Scripts | |
| The GPT-2-family runs used the `llm.c` code snapshot at Git commit | |
| `f1e2ace651495b74ae22d45d1723443fd00ecd3a`. The required CUDA trainer, headers, | |
| Makefile, dataset converter, and Hugging Face import/export utilities are preserved in | |
| `code/llmc_training/`. Exact historical launch scripts are in `training_scripts/pretrain/`. | |
| The Qwen runs used ms-swift. The 7B run-specific command, package snapshot, system record, | |
| and complete training log are in `evidence/training_runs/qwen25_coder_7b_cpt/`. The smaller | |
| May 2026 runs did not capture run-specific package snapshots; this limitation is documented | |
| in `environment/KNOWN_LIMITATIONS.md`. | |
| This S1--S4 increment contains the selected model checkpoints and their provenance | |
| catalogue. Additional S2--S4 datasets and variant-specific launch scripts remain in the | |
| local candidate staging area and are not part of this upload. | |
| Historical scripts intentionally retain the absolute paths, GPU assignments, Docker image, | |
| ports, and cache locations used at training time. Replace these machine-specific paths | |
| before rerunning. | |
| ## Training-Data Dependency | |
| CodeGPT, GPT-2, and Qwen2.5-Coder-7B were trained from the same historical source file: | |
| ```text | |
| /raid/data/weifeng/Datasets/code_messages_pre_train.jsonl | |
| 2,239,565 records; 20,021,101,140 bytes | |
| SHA-256: 74e8b9111dfca80324ade8da530858584cc3833d7650688989eedd1cb6d0f25d | |
| ``` | |
| That 20 GB raw source is not duplicated in this model-and-code increment. The GPT-2-family | |
| tokenisation and deterministic train/validation split implementation is included at | |
| `code/llmc_training/dev/data/code_messages.py`. The smaller filtered Stage-1 dataset and | |
| the complete F3 chain already remain under `data/`. | |
| ## F3 Filtering and Evaluation Evidence | |
| `data/f3/` contains the input, principal intermediate artefacts, score output, and final | |
| 207,009-record F3 training dataset. Its corresponding code is in `code/f3_filtering/`. | |
| Stage 4 used an external `Qwen3-Coder-30B-A3B-Instruct` scoring model; its saved scores and | |
| final dataset are included, while its original absolute model path must be replaced on a | |
| new system. | |
| `evidence/eval/f3_refined_9p1_summary.json` records the historical Runnable alignment | |
| evaluation: 200 tasks, five samples per task, 1,000 attempts, pass@1 of 9.1%, and pass@5 of | |
| 13.0%. This is not the later repository-level Core/Hold benchmark. | |
| ## Licensing and External Distribution | |
| See `THIRD_PARTY_NOTICES.md`. The official GPT-2 metadata declares MIT and the Qwen2.5 | |
| Coder model declares Apache-2.0. The official CodeGPT-Multilingual repository declares no | |
| model licence; its derivative checkpoint therefore requires explicit rights confirmation | |
| before distribution outside the authorised delivery channel. | |
| ## Integrity and Validation | |
| - `MANIFEST.json` maps every delivered file to its source and SHA-256 digest. | |
| - `SHA256SUMS` supports byte-level verification after transfer. | |
| - `tools/build_manifest.py` regenerates both integrity records. | |
| - `tools/smoke_test.py` loads the delivered checkpoints and checks finite, | |
| non-degenerate logits. | |
| Hard-linked files are a local storage optimisation only. Archives and Hugging Face uploads | |
| read the full file contents, so recipients receive ordinary complete files. | |