rl-training-debug-artifacts / ARTIFACT_INDEX.md
xianglinyang's picture
Add files using upload-large-folder tool
274951a verified
|
Raw History Blame Contribute Delete
2.18 kB
# Artifact index
This repository is a curated debugging handoff. It is not a claim that any
included adapter improves the base model.
## Checkpoints
| Directory | Base | Configuration | Why included |
|---|---|---|---|
| `checkpoints/4b-base-dpo-cleanprompt/step-28` | `Qwen/Qwen3.5-4B` | DPO, LR 1e-6, beta 0.1, LoRA r16/a32 | Early checkpoint with the best held-out partial-reward signal |
| `checkpoints/4b-base-dpo-cleanprompt/step-125` | `Qwen/Qwen3.5-4B` | Same run | Final checkpoint, showing late degradation |
| `checkpoints/9b-dpo-length-clean/step-8` | `prism-drift/qwen35-9b-m0-v4` | DPO on 245 length-clean pairs | Early dev improvement |
| `checkpoints/9b-dpo-length-clean/step-12` | `prism-drift/qwen35-9b-m0-v4` | Same run | Checkpoint associated with held-out reversal |
| `checkpoints/9b-grpo-probes/default-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 1e-6, beta 0.05, LoRA r16/a32 | Default probe |
| `checkpoints/9b-grpo-probes/aggressive-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 5e-6, beta 0.01, LoRA r16/a32 | LR/beta probe |
| `checkpoints/9b-grpo-probes/r32a64-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 1e-6, beta 0.05, LoRA r32/a64 | Higher-capacity probe |
Each checkpoint contains only:
- `adapter_model.safetensors`
- `adapter_config.json`
- `trainer_state.json`
- the generated checkpoint `README.md`
No optimizer, scheduler, RNG, duplicated reference adapter, tokenizer copy, or
merged full-model cache is included.
## Supporting evidence
- `evaluations/`: exact completion-level artifacts and summaries used in the
report.
- `data/`: final 4B Base DPO pairs, the 4B/9B M0 length-clean subsets, and the
final Base sweep/cohort manifests.
- `metadata/runs/`: immutable run metadata and available checkpoint-time reward
logs.
- `logs/`: selected formal-run and hyperparameter-probe logs.
- `configs/`: training configurations used during the investigation.
- `code/`: the relevant training, filtering, merge, and evaluation utilities
as they existed when this package was assembled.
- `reports/rl_training_debug_report.md`: the full narrative report.
`SHA256SUMS` covers every uploaded file other than `SHA256SUMS` itself.