Spaces:
Running
Running
|
Download README.md from HerrHruby/SFT-v4-Dataset-Viewer: direct link, hf CLI and curl.
- Browser
- Download file 1.39 kB
-
https://huggingface.co/spaces/HerrHruby/SFT-v4-Dataset-Viewer/resolve/main/README.md
- Command line
-
hf download hf://spaces/HerrHruby/SFT-v4-Dataset-Viewer/README.md
-
curl -L -o README.md https://huggingface.co/spaces/HerrHruby/SFT-v4-Dataset-Viewer/resolve/main/README.md
1.39 kB
| title: SFT v4 Dataset Viewer | |
| emoji: 🧭 | |
| colorFrom: indigo | |
| colorTo: purple | |
| sdk: static | |
| app_file: index.html | |
| pinned: false | |
| license: mit | |
| short_description: Midtrain v4 SFT traces with repaired MR reasoning | |
| # SFT v4 Dataset Viewer | |
| Thirty meta-reasoning trajectories from the midtrain **v4** corpus, sampled | |
| across the full range of depth (2–11 layers) and judge score. | |
| Each trajectory shows, layer by layer: | |
| - the **MR reasoning** — the assistant target SFT trains on, with the | |
| directions it commits to; | |
| - the **executions** that pursued those directions, and their reports; | |
| - the **summaries retained into the frontier**, which is what the next layer | |
| plans over; | |
| - the **termination decision** and **final answer**, with the judge's verdict. | |
| ## About the MR reasoning shown here | |
| Deliberation originally visited directions in F2 *rank* order. Rank encodes the | |
| accept/reject decision, so the narrative came out sorted by disposition: every | |
| kept direction before every discarded one in 96% of layers, against 40% in the | |
| order the planner actually proposed them. Training on that teaches position | |
| rather than judgement. | |
| The generator now follows F1's proposal order, and the affected records were | |
| regenerated. **This viewer shows only the repaired version** — trajectories | |
| whose MR had not been regenerated are excluded rather than displayed with the | |
| original ordering. | |