ycchen commited on
Commit
3345aa0
·
verified ·
1 Parent(s): d7b8a4b

README.md: dtype bfloat16 fix / s150 layout note

Browse files
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -28,7 +28,7 @@ continuation distillation → agentic on-policy distillation (OPD).
28
  | `step05-yarn256k-config/` | 5 | Config-only YaRN-256k variant of step 4 (same weights; YaRN factor 8→32, attention factor 1.347, max position 262,144) |
29
  | `step07-softdistill-32b-config/` | 7 | Training config of the initial soft-distilled checkpoint; identical weights are public as `soft-distill-32b-deploy` in the deploy bundle |
30
  | `step09-softdistill-v2test-32b/` | 9 | Final proof-agent soft-distilled 32B (2 epochs, final loss 0.114) — student init for OPD and target for DFlash |
31
- | `step10-opd-32b-s150/` | 10 | OPD step-150 checkpoint; the released step-200 target is `opd-32b-deploy` in the deploy bundle |
32
 
33
  All full checkpoints are consolidated bf16 safetensors with config, tokenizer, and chat template.
34
  The 65 GB checkpoints load with `trust_remote_code` where custom `olmo3_sink` modeling code is present.
 
28
  | `step05-yarn256k-config/` | 5 | Config-only YaRN-256k variant of step 4 (same weights; YaRN factor 8→32, attention factor 1.347, max position 262,144) |
29
  | `step07-softdistill-32b-config/` | 7 | Training config of the initial soft-distilled checkpoint; identical weights are public as `soft-distill-32b-deploy` in the deploy bundle |
30
  | `step09-softdistill-v2test-32b/` | 9 | Final proof-agent soft-distilled 32B (2 epochs, final loss 0.114) — student init for OPD and target for DFlash |
31
+ | `step10-opd-32b-s150/` | 10 | OPD step-150 checkpoint; the released step-200 target is `opd-32b-deploy` in the deploy bundle. NOTE: this directory uses the serving layout (legacy `rope_scaling` keys, hybrid-SWA fields, no bundled custom modeling code) — load it with SGLang/the deploy stack, not stock `transformers` |
32
 
33
  All full checkpoints are consolidated bf16 safetensors with config, tokenizer, and chat template.
34
  The 65 GB checkpoints load with `trust_remote_code` where custom `olmo3_sink` modeling code is present.