Instructions to use HilaryTorn/rl-training-debug-artifacts with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use HilaryTorn/rl-training-debug-artifacts with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
|
Download ARTIFACT_INDEX.md from HilaryTorn/rl-training-debug-artifacts: direct link, hf CLI and curl.
- Browser
- Download file 2.18 kB
-
https://huggingface.co/HilaryTorn/rl-training-debug-artifacts/resolve/main/ARTIFACT_INDEX.md
- Command line
-
hf download hf://HilaryTorn/rl-training-debug-artifacts/ARTIFACT_INDEX.md
-
curl -L -o ARTIFACT_INDEX.md https://huggingface.co/HilaryTorn/rl-training-debug-artifacts/resolve/main/ARTIFACT_INDEX.md
2.18 kB
| # Artifact index | |
| This repository is a curated debugging handoff. It is not a claim that any | |
| included adapter improves the base model. | |
| ## Checkpoints | |
| | Directory | Base | Configuration | Why included | | |
| |---|---|---|---| | |
| | `checkpoints/4b-base-dpo-cleanprompt/step-28` | `Qwen/Qwen3.5-4B` | DPO, LR 1e-6, beta 0.1, LoRA r16/a32 | Early checkpoint with the best held-out partial-reward signal | | |
| | `checkpoints/4b-base-dpo-cleanprompt/step-125` | `Qwen/Qwen3.5-4B` | Same run | Final checkpoint, showing late degradation | | |
| | `checkpoints/9b-dpo-length-clean/step-8` | `prism-drift/qwen35-9b-m0-v4` | DPO on 245 length-clean pairs | Early dev improvement | | |
| | `checkpoints/9b-dpo-length-clean/step-12` | `prism-drift/qwen35-9b-m0-v4` | Same run | Checkpoint associated with held-out reversal | | |
| | `checkpoints/9b-grpo-probes/default-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 1e-6, beta 0.05, LoRA r16/a32 | Default probe | | |
| | `checkpoints/9b-grpo-probes/aggressive-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 5e-6, beta 0.01, LoRA r16/a32 | LR/beta probe | | |
| | `checkpoints/9b-grpo-probes/r32a64-step-16` | `prism-drift/qwen35-9b-m0-v4` | LR 1e-6, beta 0.05, LoRA r32/a64 | Higher-capacity probe | | |
| Each checkpoint contains only: | |
| - `adapter_model.safetensors` | |
| - `adapter_config.json` | |
| - `trainer_state.json` | |
| - the generated checkpoint `README.md` | |
| No optimizer, scheduler, RNG, duplicated reference adapter, tokenizer copy, or | |
| merged full-model cache is included. | |
| ## Supporting evidence | |
| - `evaluations/`: exact completion-level artifacts and summaries used in the | |
| report. | |
| - `data/`: final 4B Base DPO pairs, the 4B/9B M0 length-clean subsets, and the | |
| final Base sweep/cohort manifests. | |
| - `metadata/runs/`: immutable run metadata and available checkpoint-time reward | |
| logs. | |
| - `logs/`: selected formal-run and hyperparameter-probe logs. | |
| - `configs/`: training configurations used during the investigation. | |
| - `code/`: the relevant training, filtering, merge, and evaluation utilities | |
| as they existed when this package was assembled. | |
| - `reports/rl_training_debug_report.md`: the full narrative report. | |
| `SHA256SUMS` covers every uploaded file other than `SHA256SUMS` itself. | |