Spaces:
Running
Download README.md from HerrHruby/SFT-v4-Dataset-Viewer: direct link, hf CLI and curl.
- Browser
- Download file 1.39 kB
-
https://huggingface.co/spaces/HerrHruby/SFT-v4-Dataset-Viewer/resolve/main/README.md
- Command line
-
hf download hf://spaces/HerrHruby/SFT-v4-Dataset-Viewer/README.md
-
curl -L -o README.md https://huggingface.co/spaces/HerrHruby/SFT-v4-Dataset-Viewer/resolve/main/README.md
title: SFT v4 Dataset Viewer
emoji: 🧭
colorFrom: indigo
colorTo: purple
sdk: static
app_file: index.html
pinned: false
license: mit
short_description: Midtrain v4 SFT traces with repaired MR reasoning
SFT v4 Dataset Viewer
Thirty meta-reasoning trajectories from the midtrain v4 corpus, sampled across the full range of depth (2–11 layers) and judge score.
Each trajectory shows, layer by layer:
- the MR reasoning — the assistant target SFT trains on, with the directions it commits to;
- the executions that pursued those directions, and their reports;
- the summaries retained into the frontier, which is what the next layer plans over;
- the termination decision and final answer, with the judge's verdict.
About the MR reasoning shown here
Deliberation originally visited directions in F2 rank order. Rank encodes the accept/reject decision, so the narrative came out sorted by disposition: every kept direction before every discarded one in 96% of layers, against 40% in the order the planner actually proposed them. Training on that teaches position rather than judgement.
The generator now follows F1's proposal order, and the affected records were regenerated. This viewer shows only the repaired version — trajectories whose MR had not been regenerated are excluded rather than displayed with the original ordering.