|
Download README.md from SlayerLab/Slayer149-checkpoints: direct link, hf CLI and curl.
- Browser
- Download file 1.56 kB
-
https://huggingface.co/SlayerLab/Slayer149-checkpoints/resolve/main/README.md
- Command line
-
hf download hf://SlayerLab/Slayer149-checkpoints/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/Slayer149-checkpoints/resolve/main/README.md
1.56 kB
| language: en | |
| license: apache-2.0 | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| tags: | |
| - experimental | |
| - checkpoints | |
| # Slayer149 training checkpoints | |
| Automatic recovery archive for controlled data-continuation experiments from | |
| [SlayerLab/Slayer149](https://huggingface.co/SlayerLab/Slayer149). | |
| 148,910,738 parameters. Experimental checkpoints are not promoted releases or | |
| verified top-three leaderboard entries. The original release remains unchanged. | |
| `baseline/training-state.pt` and each `runs/<arm-seed>/checkpoint-<step>/training-state.pt` | |
| contain weights, optimizer state, step, training configuration and provenance. | |
| `extension/` contains the longer continuation only if the development selection | |
| criteria pass. Steps are local to each training phase, not total lineage tokens. | |
| Safetensors inference exports are included for the final pilot/extension checkpoints. | |
| These use the accompanying custom PyTorch loader, not Transformers AutoModel. | |
| `reports/` contains development measurements and experiment status. Development | |
| proxies are not GLINT leaderboard results. Full GLINT confirmation, when available, | |
| is explicitly named. Scores and ranking claims must be read with their protocol. | |
| No training corpus or authentication credentials are uploaded. | |
| Restore a trusted `training-state.pt` as `latest.pt` in a new run directory, | |
| restore its matching config and exact tokenized data, then resume the trainer. | |
| Source/tokenizer revisions and data hashes are recorded in the manifests; this | |
| archive does not by itself contain the training datasets. | |