--- title: SFT v4 Dataset Viewer emoji: 🧭 colorFrom: indigo colorTo: purple sdk: static app_file: index.html pinned: false license: mit short_description: Midtrain v4 SFT traces with repaired MR reasoning --- # SFT v4 Dataset Viewer Thirty meta-reasoning trajectories from the midtrain **v4** corpus, sampled across the full range of depth (2–11 layers) and judge score. Each trajectory shows, layer by layer: - the **MR reasoning** — the assistant target SFT trains on, with the directions it commits to; - the **executions** that pursued those directions, and their reports; - the **summaries retained into the frontier**, which is what the next layer plans over; - the **termination decision** and **final answer**, with the judge's verdict. ## About the MR reasoning shown here Deliberation originally visited directions in F2 *rank* order. Rank encodes the accept/reject decision, so the narrative came out sorted by disposition: every kept direction before every discarded one in 96% of layers, against 40% in the order the planner actually proposed them. Training on that teaches position rather than judgement. The generator now follows F1's proposal order, and the affected records were regenerated. **This viewer shows only the repaired version** — trajectories whose MR had not been regenerated are excluded rather than displayed with the original ordering.