Instructions to use HilaryTorn/rl-training-debug-artifacts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HilaryTorn/rl-training-debug-artifacts with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download ARTIFACT_INDEX.md from HilaryTorn/rl-training-debug-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 2.18 kB
-
https://huggingface.co/HilaryTorn/rl-training-debug-artifacts/resolve/main/ARTIFACT_INDEX.md
- Command line
-
hf download hf://HilaryTorn/rl-training-debug-artifacts/ARTIFACT_INDEX.md
-
curl -L -o ARTIFACT_INDEX.md https://huggingface.co/HilaryTorn/rl-training-debug-artifacts/resolve/main/ARTIFACT_INDEX.md
2.18 kB
Artifact index
This repository is a curated debugging handoff. It is not a claim that any included adapter improves the base model.
Checkpoints
| Directory | Base | Configuration | Why included |
|---|---|---|---|
checkpoints/4b-base-dpo-cleanprompt/step-28 |
Qwen/Qwen3.5-4B |
DPO, LR 1e-6, beta 0.1, LoRA r16/a32 | Early checkpoint with the best held-out partial-reward signal |
checkpoints/4b-base-dpo-cleanprompt/step-125 |
Qwen/Qwen3.5-4B |
Same run | Final checkpoint, showing late degradation |
checkpoints/9b-dpo-length-clean/step-8 |
prism-drift/qwen35-9b-m0-v4 |
DPO on 245 length-clean pairs | Early dev improvement |
checkpoints/9b-dpo-length-clean/step-12 |
prism-drift/qwen35-9b-m0-v4 |
Same run | Checkpoint associated with held-out reversal |
checkpoints/9b-grpo-probes/default-step-16 |
prism-drift/qwen35-9b-m0-v4 |
LR 1e-6, beta 0.05, LoRA r16/a32 | Default probe |
checkpoints/9b-grpo-probes/aggressive-step-16 |
prism-drift/qwen35-9b-m0-v4 |
LR 5e-6, beta 0.01, LoRA r16/a32 | LR/beta probe |
checkpoints/9b-grpo-probes/r32a64-step-16 |
prism-drift/qwen35-9b-m0-v4 |
LR 1e-6, beta 0.05, LoRA r32/a64 | Higher-capacity probe |
Each checkpoint contains only:
adapter_model.safetensorsadapter_config.jsontrainer_state.json- the generated checkpoint
README.md
No optimizer, scheduler, RNG, duplicated reference adapter, tokenizer copy, or merged full-model cache is included.
Supporting evidence
evaluations/: exact completion-level artifacts and summaries used in the report.data/: final 4B Base DPO pairs, the 4B/9B M0 length-clean subsets, and the final Base sweep/cohort manifests.metadata/runs/: immutable run metadata and available checkpoint-time reward logs.logs/: selected formal-run and hyperparameter-probe logs.configs/: training configurations used during the investigation.code/: the relevant training, filtering, merge, and evaluation utilities as they existed when this package was assembled.reports/rl_training_debug_report.md: the full narrative report.
SHA256SUMS covers every uploaded file other than SHA256SUMS itself.