File size: 9,809 Bytes
b7e9b58 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 | # Reproducing the recorded analysis and checkpoint evaluation
This guide is new release documentation. The copied source files, models, logs, reports, and original job files remain unchanged. Assembly checks were performed; this packaging task did not run new Unity evaluations or retrain models.
## 1. Create a separate working copy
Use Python 3.10 or newer for the preparation helper. Choose a **new, short directory outside Upload**, such as `D:/SGRepro`, to avoid long Windows package paths. The helper refuses an existing destination and does not delete or overwrite it.
From the Upload directory:
```powershell
python 07_Reproduction/prepare_workspace.py --destination D:/SGRepro --dry-run
python 07_Reproduction/prepare_workspace.py --destination D:/SGRepro
```
The helper only copies files and prepares new configuration files; it never starts Unity, training, or evaluation. The resulting layout is:
```text
D:/SGRepro/
RevisionEval/ # Unity project, including model assets
Document/
ShootingGame_IDE_Handoff/ # validators and original model/dataset modules
RevisionEvaluation_Local_20260910/
RevisionAdditional_20260911/
Tools/
Results/ # original recorded results
ArchivedReports/ # original evidence, unchanged
OriginalJobs/ # original job bytes, unchanged
Jobs/ # new copies with rebased output/report paths
ReproducedResults/ # fresh evaluation output destination
Reports/ # newly generated reports
Logs/
Verification/
```
The original jobs contain the authors' absolute paths. The helper changes **only `output` and `report` in the newly generated job copies**; seeds, episode counts, policy, model paths, target counts, RTGs, and diagnostic/timing flags are preserved. Archived results are kept separately from fresh evaluation outputs. The helper restores the directory layout for the direct commands documented below; moving individual scripts alone is insufficient. `run_phase.py` and `run_additional.py` are preserved historical orchestrators, not the entry points for this new workspace: they expect earlier phase-state reports and the original result-directory conventions.
## 2. Validate the recorded episode files
From `D:/SGRepro`, use Python to run these read-only validators:
```powershell
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionEvaluation_Local_20260910/Results/full --expected-episodes 50 --expected-models 19 --expected-targets 10 15 20
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionAdditional_20260911/Results/heldout --expected-episodes 200 --expected-models 3 --expected-targets 10 15 20
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionAdditional_20260911/Results/rtg --expected-episodes 50 --expected-models 2 --expected-targets 20 --allow-rtg-sensitivity
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionAdditional_20260911/Results/bc-diagnostic/bc-normal --expected-episodes 50 --expected-models 1 --expected-targets 20
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionAdditional_20260911/Results/bc-diagnostic/bc-legacy-rtg --expected-episodes 50 --expected-models 1 --expected-targets 20 --allow-bc-diagnostic
python Document/RevisionAdditional_20260911/Tools/validate_additional.py Document/RevisionAdditional_20260911/Results/timing --expected-episodes 20 --expected-models 3 --expected-targets 20
```
The normal and legacy BC diagnostic folders must be checked separately. The diagnostic legacy condition is not a valid constant-zero BC evaluation and must not be pooled with the main comparison.
## 3. Recompute the main statistical reports
Use an isolated Python environment with NumPy **1.26.4** and SciPy **1.13.1** to match the recorded main analysis versions. For example, Python 3.12 supports these versions. These analysis requirements do not reconstruct the original training environment.
```powershell
python -m pip install numpy==1.26.4 scipy==1.13.1
python Document/RevisionAdditional_20260911/Tools/analyze_statistics.py Document/RevisionEvaluation_Local_20260910/Results/full existing
python Document/RevisionAdditional_20260911/Tools/analyze_statistics.py Document/RevisionAdditional_20260911/Results/heldout heldout
python Document/RevisionAdditional_20260911/Tools/analyze_diagnostics.py
python Document/RevisionAdditional_20260911/Tools/analyze_reward_components.py
```
Keep the analysis names **`existing` and `heldout`** unchanged: they participate in deterministic resampling seed labels. New reports are written to the working copy's `Reports` directory; archived reports remain in `ArchivedReports`.
Success uses exact binomial intervals and exact McNemar comparisons; continuous outcomes use paired sign permutation tests and bootstrap intervals. Holm correction is applied to four separate exploratory families: 126 full-comparison tests, 54 new-seed tests, 72 RTG contrasts, and 6 BC-diagnostic contrasts. Intervals are marginal, not simultaneous. These analyses are conditional on fixed checkpoints and were not preregistered.
The commands above analyze the **archived `Results`**. They do not automatically analyze fresh `ReproducedResults`; select fresh inputs explicitly for a new analysis and retain an unambiguous output record.
`report_additional.py` is preserved as an auxiliary historical report generator, not one of the direct reproduction commands above. It additionally requires Matplotlib and expects `Reports/latency-benchmark.json` for its standalone-latency section; the example's newly generated `release-latency.json` is not automatically substituted. Some auxiliary reward-bootstrap seeds use the absolute raw-file path. Moving the files can therefore change those auxiliary Monte Carlo interval endpoints. The main `existing-statistics.json` and `heldout-statistics.json` procedures above use their preserved analysis labels. Compare scientific values separately from recorded absolute provenance paths, which necessarily differ in a new workspace.
## 4. Run the Unity checkpoint evaluations
Install Unity **6000.0.62f1** with a working local license. Open `D:/SGRepro/RevisionEval` to let Unity restore packages from the included manifest and lockfile. Recorded package versions are Inference Engine **2.4.1** and ML-Agents **4.0.0**. The tested platform is Windows with Direct3D11 and a GPU; do not use `-nographics` for GPU parity/evaluation.
The original launchers expect the standard Unity Hub executable path. If your installation differs, set the `editor` path in the **working copy** of `Tools/Run-Unity.ps1`, or invoke the same `-executeMethod` and arguments with your Unity executable. Do not edit the release originals.
Example from the working-copy root:
```powershell
powershell -NoProfile -File Document/RevisionEvaluation_Local_20260910/Tools/Run-Unity.ps1 -Mode Verify -RunName release-verify
powershell -NoProfile -File Document/RevisionEvaluation_Local_20260910/Tools/Run-Unity.ps1 -Mode Run -JobPath Document/RevisionEvaluation_Local_20260910/Jobs/full-ppo-10.json -RunName release-full-ppo-10
```
Run each desired job once with its own run name. The complete full comparison has six jobs: `full-ppo-{10,15,20}` and `full-transformers-{10,15,20}`. The additional comparison has six `heldout-*` jobs; the RTG sweep has twelve `rtg-000` through `rtg-110` jobs plus `rtg-bc-reference`; the BC diagnostic has `bc-normal` and `bc-legacy-rtg`; Play-mode timing has `timing-dt-bc` and `timing-ppo`. Extra smoke/repeat jobs are initialization/reproducibility checks, not part of the reported performance cohorts.
Additional jobs use the corresponding launcher:
```powershell
powershell -NoProfile -File Document/RevisionAdditional_20260911/Tools/Run-Unity.ps1 -Mode Run -JobPath Document/RevisionAdditional_20260911/Jobs/heldout-10-dt-bc.json -RunName release-heldout-10-dt-bc
powershell -NoProfile -File Document/RevisionAdditional_20260911/Tools/Run-Unity.ps1 -Mode Verify -RunName release-additional-parity
powershell -NoProfile -File Document/RevisionAdditional_20260911/Tools/Run-Unity.ps1 -Mode Latency -RunName release-latency
```
PPO remains stochastic (`DeterministicInference=false`), using the original CPU/default backend; DT/BC use GPUCompute. Environment seeds and initial states can match while PPO actions and episode metrics vary. Timing results depend on the hardware and backend and are not rendered deployment FPS measurements. Standalone latency records 2,000 calls after 200 warmups; Play-mode timing is a separate design.
The latest expanded parity evidence is `additional-verification-retry1.json` and `pytorch-additional-parity.json`, alongside fixtures and reference-weight hashes. They cover 96 actual input cases for BC E1/E2/E3 and DT_S_100 E3. Same-archive pairing of other `.pth` and ONNX files is provenance evidence, not a claim of newly executed numerical parity for every checkpoint.
## 5. Training-data limits
The C archive is supplied as the original ZIP. Review its inventory before extracting it into a training workspace. The original notebooks and training scripts preserve their historical file-selection rules; do not assume that every archived C trajectory was used by every model.
The exact S partitions, their selection manifest, and a complete original training environment lock were not recovered. Training the S-only DT/BC and S-to-C models from scratch cannot yet be fully reproduced from this release. Historical raw-log candidates and other dated C collections have not been substituted for the missing training input.
|