# Decision Transformer / Unity evaluation artifacts This release folder assembles original research files for the repository-information request (Reviewer Minor 3): source code, data, trained weights, ONNX exports, and evaluation instructions. The files also expose the corrected episode results and statistical procedures relevant to Major 7–10 and 14. The code, models, raw evaluation JSON, reports, and original notebooks are copied without rewriting their contents. New release guides, file manifests, and the workspace-preparation helper are identified separately. This folder is a local publication package prepared on 2026-09-12; preparing it does not publish it to Hugging Face. ## Contents | Folder | Contents | |---|---| | `01_Source_Code/Unity_Evaluation` | Corrected Unity project source: Assets, `.meta` GUIDs, scenes, referenced model assets, Packages and ProjectSettings. Unity caches and old result dumps are omitted. | | `01_Source_Code/Python_Training` | Five original DT/BC dataset, model, training, and fine-tuning modules. | | `01_Source_Code/Notebooks` | Original training and historical analysis notebooks. Their historical result claims are not the corrected results. | | `01_Source_Code/Python_Evaluation` | Original evaluation launchers, validators, and analysis scripts. | | `02_Training_Data` | Availability statement for the missing, exact training-phase S data. No replacement dataset is labeled as S. | | `03_Expert_Data` | Original post-training C archive and its provenance; see the folder README for the exact contents and limits. | | `04_Model_Weights/PyTorch` | DT and BC PyTorch weights with original names. PPO's evaluated trained policy is provided as ONNX; a matching native V12 checkpoint was not located. | | `05_ONNX_Models` | The exact 19 evaluated exports: PPO 1, DT 15, BC 3. | | `06_Evaluation_Results` | Corrected raw results and original statistical, execution, latency, parity, and provenance reports. | | `07_Reproduction` | Original jobs and fixtures, original work notes, and instructions for a separate reproduction workspace. | | `08_Manifests` | File hashes, provenance, and package verification. | The Unity source tree includes older model assets referenced by the original project. The authoritative **19-model evaluated set** is `05_ONNX_Models` and its manifest; do not count every asset in the Unity source tree as an evaluated checkpoint. Copies of the evaluated ONNX files inside the Unity project preserve the original asset paths and GUIDs. ## Corrected evaluation cohorts | Cohort | Files | Episodes | Seeds | Interpretation | |---|---:|---:|---|---| | Full comparison, `RevisionEvaluation_Local_20260910/Results/full` | 57 | 2,850 | 42–91 | 19 checkpoints × 3 target counts × 50 episodes | | New-seed comparison, `RevisionAdditional_20260911/Results/heldout` | 9 | 1,800 | 1000–1199 | DT_S_100 E3, BC E3, PPO × 3 target counts × 200 episodes | | RTG sensitivity, `.../Results/rtg` | 13 | 650 | 2000–2049 | DT RTG 0–110 in increments of 10, plus a separate constant-zero BC reference | | BC update diagnostic, `.../Results/bc-diagnostic` | 2 | 100 | 42–91 | Correct constant-zero BC versus a diagnostic legacy decrement rule | | Play-mode timing, `.../Results/timing` | 3 | 60 | 3000–3019 | Timing episodes, separate from performance cohorts | These cohorts must not be pooled as independent evidence. The BC-normal diagnostic repeats the corresponding full-comparison condition. E1/E2/E3 are checkpoints along a training run, not independent training seeds. DT starts at RTG **35/55/70** for **10/15/20 targets** and subtracts observed rewards. Correct BC keeps **every RTG at zero**. The exploratory RTG sweep did not select or change these retained main settings. Original selection timing relative to the historical evaluation is not documented. ## Reproduction and limits Start with [the reproduction guide](07_Reproduction/REPRODUCE.md). It preserves the released originals, reconstructs the directory layout for the direct commands documented in the guide, and creates new job copies with fresh output paths. Do not execute the original absolute-path job files directly. Historical batch orchestrators need their old phase-state reports and are preserved for reference rather than used by this guide. The tested environment was Unity **6000.0.62f1**, Inference Engine **2.4.1**, ML-Agents **4.0.0**, Windows/Direct3D11. DT/BC used GPUCompute; PPO used the ML-Agents CPU/default backend with stochastic inference. Matched environmental seeds do not make PPO action sampling bitwise deterministic. The main statistical reports used **NumPy 1.26.4** and **SciPy 1.13.1**. The later four-checkpoint parity report used **PyTorch 2.6.0+cu124**, **ONNX 1.19.1**, and **ONNX Runtime 1.27.0**. These are analysis/verification environments, not a recovered lockfile for the original training environment. **Exact S training trajectories and the final S `.pkl` partitions remain unavailable.** Historical raw-log candidates were located, but their identity with the training input was not established, so they are not substituted here. Evaluation of the supplied checkpoints and analysis of recorded results are supported; complete from-scratch reproduction of S-dependent training remains incomplete. No new training or evaluation was performed to assemble this package. The pinned source for recovered public originals is [repository revision b68573995aea82d97a599ee7ea0c9b2ea8d88f8b](https://huggingface.co/code3939/DecisionTransformer-Unity-Sim/tree/b68573995aea82d97a599ee7ea0c9b2ea8d88f8b). Historical notebooks and archived notes may contain earlier paths or interpretations; use the cohort definitions above and the corrected statistical reports for this release.