rl-training-debug-artifacts / ARTIFACT_INDEX.md
xianglinyang's picture
Add files using upload-large-folder tool
274951a verified
|
Raw History Blame Contribute Delete
2.18 kB

Artifact index

This repository is a curated debugging handoff. It is not a claim that any included adapter improves the base model.

Checkpoints

Directory Base Configuration Why included
checkpoints/4b-base-dpo-cleanprompt/step-28 Qwen/Qwen3.5-4B DPO, LR 1e-6, beta 0.1, LoRA r16/a32 Early checkpoint with the best held-out partial-reward signal
checkpoints/4b-base-dpo-cleanprompt/step-125 Qwen/Qwen3.5-4B Same run Final checkpoint, showing late degradation
checkpoints/9b-dpo-length-clean/step-8 prism-drift/qwen35-9b-m0-v4 DPO on 245 length-clean pairs Early dev improvement
checkpoints/9b-dpo-length-clean/step-12 prism-drift/qwen35-9b-m0-v4 Same run Checkpoint associated with held-out reversal
checkpoints/9b-grpo-probes/default-step-16 prism-drift/qwen35-9b-m0-v4 LR 1e-6, beta 0.05, LoRA r16/a32 Default probe
checkpoints/9b-grpo-probes/aggressive-step-16 prism-drift/qwen35-9b-m0-v4 LR 5e-6, beta 0.01, LoRA r16/a32 LR/beta probe
checkpoints/9b-grpo-probes/r32a64-step-16 prism-drift/qwen35-9b-m0-v4 LR 1e-6, beta 0.05, LoRA r32/a64 Higher-capacity probe

Each checkpoint contains only:

  • adapter_model.safetensors
  • adapter_config.json
  • trainer_state.json
  • the generated checkpoint README.md

No optimizer, scheduler, RNG, duplicated reference adapter, tokenizer copy, or merged full-model cache is included.

Supporting evidence

  • evaluations/: exact completion-level artifacts and summaries used in the report.
  • data/: final 4B Base DPO pairs, the 4B/9B M0 length-clean subsets, and the final Base sweep/cohort manifests.
  • metadata/runs/: immutable run metadata and available checkpoint-time reward logs.
  • logs/: selected formal-run and hyperparameter-probe logs.
  • configs/: training configurations used during the investigation.
  • code/: the relevant training, filtering, merge, and evaluation utilities as they existed when this package was assembled.
  • reports/rl_training_debug_report.md: the full narrative report.

SHA256SUMS covers every uploaded file other than SHA256SUMS itself.