memory-models / README.md
Caesarrr's picture
Consolidate complete memory mechanism model archive
2c0261c
|
Raw History Blame Contribute Delete
2.61 kB

Memory mechanism experiment models

Research archive for Flappy Bird, Demon Attack and Deadly Corridor. The archive preserves 20 QwenOFT runs, three WanOFT context5 runs, and three rollout teachers. It does not establish a new benchmark result.

Family Archived conditions Runs Retained training step
QwenOFT / Flappy L0 single baseline; L2 single, multi, stitch, KV and flex KV; L3 single, multi, stitch and flex KV 10 4000
QwenOFT / Demon Attack L0 single baseline; L6 single, multi, stitch and flex KV 5 4000
QwenOFT / Deadly Corridor L0 single baseline; L6 single, multi, stitch and flex KV 5 4000
WanOFT context5 Flappy L3; Demon Attack L6; Deadly Corridor L6 3 2000; 5000; 1000
Sample Factory rollout teachers Flappy L3; Demon Attack L6; Deadly Corridor L6 3 48840; 48880; 48836

L denotes latency in raw environment frames. Input modes, timing, loss contracts and normalization statistics remain in each run's config. KV training sequence length and inference window are separate settings.

Download and provenance

run_index.json lists exact checkpoint paths, original source identities and fixed data references. Configs and normalization files remain beside the weights in the original run directories.

Training data and original logs/evaluation evidence are in memory-data@49da56bd92b9842fb457418aba760a1b03379258. All seven training configs retain their original paths. QwenOFT L0, Flappy L2, and the QwenOFT/WanOFT L3/L6 runs each have their corresponding dataset listed in the index.

Weight preservation

The 19 independent QwenOFT weights and three WanOFT weights retain their source content identities. The old KV memory ZeRO state contains a complete model and was exported as flappy_fix_latency_2_200ep_7k2steps_kv_memory/checkpoints/steps_4000_pytorch_model.pt. All 730 model tensors, their dtypes and PyTorch module metadata were checked before native loading.

The three teacher files retain all model parameters, normalization statistics, module metadata, train steps and environment steps. Optimizer and other training recovery state were removed. All 26 retained weights passed actual download and the corresponding native loader.

Historical intermediate and .pre_bce evaluations remain evidence in the data archive. Their presence does not imply that missing intermediate weights are available or that historical results are accepted benchmark results.