YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Memory mechanism experiment models
Research archive for Flappy Bird, Demon Attack and Deadly Corridor. The archive preserves 20 QwenOFT runs, three WanOFT context5 runs, and three rollout teachers. It does not establish a new benchmark result.
| Family | Archived conditions | Runs | Retained training step |
|---|---|---|---|
| QwenOFT / Flappy | L0 single baseline; L2 single, multi, stitch, KV and flex KV; L3 single, multi, stitch and flex KV | 10 | 4000 |
| QwenOFT / Demon Attack | L0 single baseline; L6 single, multi, stitch and flex KV | 5 | 4000 |
| QwenOFT / Deadly Corridor | L0 single baseline; L6 single, multi, stitch and flex KV | 5 | 4000 |
| WanOFT context5 | Flappy L3; Demon Attack L6; Deadly Corridor L6 | 3 | 2000; 5000; 1000 |
| Sample Factory rollout teachers | Flappy L3; Demon Attack L6; Deadly Corridor L6 | 3 | 48840; 48880; 48836 |
L denotes latency in raw environment frames. Input modes, timing, loss contracts and normalization statistics remain in each run's config. KV training sequence length and inference window are separate settings.
Download and provenance
run_index.json lists exact checkpoint paths, original source identities and fixed data references. Configs and normalization files remain beside the weights in the original run directories.
Training data and original logs/evaluation evidence are in memory-data@49da56bd92b9842fb457418aba760a1b03379258. All seven training configs retain their original paths. QwenOFT L0, Flappy L2, and the QwenOFT/WanOFT L3/L6 runs each have their corresponding dataset listed in the index.
Weight preservation
The 19 independent QwenOFT weights and three WanOFT weights retain their source content identities. The old KV memory ZeRO state contains a complete model and was exported as flappy_fix_latency_2_200ep_7k2steps_kv_memory/checkpoints/steps_4000_pytorch_model.pt. All 730 model tensors, their dtypes and PyTorch module metadata were checked before native loading.
The three teacher files retain all model parameters, normalization statistics, module metadata, train steps and environment steps. Optimizer and other training recovery state were removed. All 26 retained weights passed actual download and the corresponding native loader.
Historical intermediate and .pre_bce evaluations remain evidence in the data archive. Their presence does not imply that missing intermediate weights are available or that historical results are accepted benchmark results.