LeWAM OGBench puzzle (3x3)
Goal-reaching checkpoint in the LeWAM v1 format. Files are stored at the repository root:
lewam_best.pt
lewam_config.json
Download both files into $STABLEWM_HOME/checkpoints/<run_name>/, then evaluate with the v1 code:
python scripts/eval_lewam.py --config-name puzzle policy=<run_name>
The checkpoint was trained with lewam-jointflow@2d52888 and converted to the v1 schema.
Every retained tensor is exactly equal to the original after renaming; strict model loading,
attention masks, encoder output, supplied-goal action prediction and dynamics prediction passed.
Only the obsolete, all-zero null_goal tensor was removed; the original goal-dropout setting was zero.
No optimizer state is included.
Original checkpoint and historical performance: source repository. Drawer uses goal offset 50 and budget 100; the other tasks use offset 25 and budget 50.
SHA256 of lewam_best.pt: da432b0ac9d9d0c55207ed0db2323bf01be96d9ccfecd7026a7362ce7fef788d.
Local v1 reactive evaluation (50 episodes per seed, seeds 42/0/1): 78 / 84 / 70% success. The historical reference is 78 / 84 / 68%; the seed-1 difference is one episode and its cause has not been established.