YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- CommitWAM-v0.1 corrected implementation report
- 1. Experimental scope and frozen contract
- 2. Protected old run
- 3. Corrected dataset construction
- 4. Feature/cache contract and causality
- 5. Timing-shortcut controls
- 6. Model, objective, and batching fixes
- 7. Tests and preflight
- 8. Final-code GPU smoke
- 9. Production training result
- 10. Post-training feature diagnostics
- 11. Reproducibility and explicit confirmations
- 1. Experimental scope and frozen contract
CommitWAM-v0.1 corrected implementation report
Date: 2026-09-14
Branch: CommitWAM-v01
Training mode: one GPU in the existing interactive Slurm allocation; no sbatch submission.
This report records the corrected, controller-cadence-aligned BCE experiment. It supersedes the earlier report that described U/U+4 positives and the old 811950 launch. The old run is preserved as a diagnostic only.
1. Experimental scope and frozen contract
The experiment tests one narrow question:
Given that frozen Qwen has proposed a distinct next semantic phase, does the causal execution history support committing away from the active phase?
The only trainable component is CommitReadinessHead (4,416,513 parameters). The following remain frozen and unchanged:
Qwen/Qwen3-VL-4B-Instructand the existing simple-subgoal adapter/checkpoint;- the FastWAM executor and Memory Expert;
- SigLIP and the text encoder;
simple_subgoalprompt/history semantics;- 20 predicted actions, 16 executed actions, and 512-token FastWAM memory.
No candidate text is supplied to the head. No human labels, candidate-validity/failure/progress head, WAIT/HOLD action, asynchronous planner, Qwen/FastWAM/Memory-Expert/SigLIP/text training, or test-set tuning was used.
The implementation is under src/fastwam/commitwam/; the production configuration is v01_grid_bce.yaml. The runtime controller remains deferred to the evaluation server; threshold selection must use validation rollouts only.
2. Protected old run
The completed job-811950 directory was not opened for writing, resumed, renamed, deleted, or used for initialization:
/groups/yshang/an221229/checkpoints/FastWAM_Memory/runs/commitwam_v01/bce_seed7
It is retained as the “clock-confounded diagnostic run.” Recorded values before the corrected data build were:
- 20 epochs and 5,080 optimizer updates;
- 4,416,513 trainable parameters;
- best reported AUPRC 0.9997215696542082 (epoch 12, step 3048);
- best reported NLL 0.0207177230782142 (epoch 10, step 2540);
- pairwise accuracy 1.0.
Read-only identity records are retained in the audit/review history: old resolved-config hash 4dd3303aa6e629644b056ecb15ee872139243ef7b180a1200b1347d2adb3227e, old implementation hash 12e6263e48de461870be48e8c268941694f84e33712020993ba647cfaa4ec822, old anchors hash d594a378e3cbe41fbd2ae6c4730438f3d25a50cd6bcdeb33b2b3f9313bf36f6f, old transitions hash 07629f46e426c358e43442d53a2ad8e9c823b3af1a27f3f95f52a44edf898d5a, and old run manifest hash e3b7662d3737f3708779fb3de21b37bf8ac54ecea3f9346a87f08334f5b2082b. The old data-sidecar hashes were recorded before creating the isolated corrected sidecar.
The protected checkpoint hashes were rechecked after all diagnostics: last.pt 39b0004c5438911e37d5f0fac498094acdbedb0e091deea0c1c3d3415dd28316, best_auprc.pt 5c246c4b5058f206c8bb1d55b67ef464e15af94c74ca4de1412e13d225b0727c, and best_nll.pt f806d82201dc6b40a3c253b6bde22db2966a879970e191f6ae0eff17abc7aa22. The old metrics hash is af161cee2177bb88b7f36b74ce9e8c8a593b62961ad009c2e5b96654b5b25b09.
The old model was evaluated read-only after corrected training. Its feature-ablation output is old_job_811950_feature_ablations.json, outside the protected run directory.
3. Corrected dataset construction
The corrected sidecar is isolated at:
/groups/yshang/an221229/data/FastWAM_Memory/commitwam/v01_grid
The source sidecar at .../commitwam/v01 remains unchanged. The corrected builder uses generic interval fields (transition_lower_bound, transition_upper_bound, transition_source) and retains native-only transitions for diagnostics, but trains only usable distinct-text transitions.
Transition audit
- 1,600 episodes: 16 task types × 100;
- 1,440 train episodes and 160 offline-dev episodes;
- 7,456 total transition rows, including native-only rows;
- 3,957 usable distinct-text transitions;
- 3,957 dual-text/native-aligned transitions; 3,499 native-only diagnostic transitions;
- online boundary earlier than offline for 3,931 transitions, equal for 26, later for 0;
- interval width mean 16.576, median 14, p90 34, maximum 46;
- native alignment rate 1.0.
Anchor labels
The deployment cadence is C=16. For each usable transition, the primary positive is the first controller anchor at or after the upper bound:
t_pos = ((U + C - 1) // C) * C
U and U+4 are not BCE positives in this run. Delayed-active proxies and a second +16 positive are disabled. Safe negatives are natural 16-step anchors satisfying t <= L-16. Natural anchors in the uncertain interval are retained as ambiguous_interval, but have target_mask=false and loss_weight=0.
Later active starts use the same grid ceiling of the predecessor upper bound; the first phase starts at zero. Incompatible cases are skipped and counted rather than repaired with future metadata.
| Quantity | Count |
|---|---|
| Stored anchors | 21,996 |
| Safe negatives | 10,171 |
| Primary grid positives | 3,906 |
| Ambiguous masked anchors | 7,919 |
| Trainable anchors | 14,077 |
| Trainable train rows | 12,613 (9,115 negative + 3,498 positive) |
| Trainable offline-dev rows | 1,464 (1,056 negative + 408 positive) |
For direct comparison, the protected clock-confounded table had 26,000 stored anchors and 18,081 trainable rows (16,197 train + 1,884 offline-dev), including 3,957 U positives and 3,953 U+4 positives. The corrected table therefore removes 4,004 stored rows and 4,004 trainable rows while retaining only one controller-grid primary positive per usable transition.
Schedule/audit counts: 3,957 usable transitions; 3,691 with a safe negative; 3,906 with a primary grid positive; 43 positives skipped because they crossed the next transition; 8 active-start incompatibilities (7 after phase, 1 minimum dwell); 3,947 transitions with at least one trainable example; 10 without any trainable example. Every trainable positive and negative is on the 16-step grid (alignment rate 100%).
Important hashes:
episode_map.parquet:9f80bdc93c847af10470d9d724e4b2858f7b31c3736fdba6509f7e615ec01ac5;splits.json:79218130aced16782221f0dd007cfc960652bfa9d77dcafa0e8eabeef7906b41;transitions.parquet:07629f46e426c358e43442d53a2ad8e9c823b3af1a27f3f95f52a44edf898d5a;- corrected
anchors.parquet:7726b673e08c5cb9f54a66703c8e9fb4baa9410db869ec7125392038aceeae42; normalization.json:81d4d6e892ff6161c0868e33dab5fc5a23c2e5d0ac4388c12924308a19f11478;- audit identity:
7f39efdea41d348599133f4791873b6ef9d8501bb347b8ac30cece7e1c618e0a.
The complete stored audit is full_audit.json and reports passed=true with no errors.
4. Feature/cache contract and causality
The model-facing allow-list remains ten fields: causal front features/history and ages, active-start front/state, recent front/state, current wrist, masked task-plus-active-subgoal text, and labels/metadata kept outside the model feature dictionary. Visual input is 608 SigLIP tokens of width 1,152. Active-start and recent states have separate role embeddings. No candidate text, future observation/action, transition bound, success flag, or label-source field is passed as a feature.
The corrected wrist cache was rebuilt for the new anchors (1,407 episode files, 12,002 frames). The expensive front cache was reused only after exact manifest verification; its manifest hash is 961d7c120f05f69e6e33aefa859ae80444b9d13449d24e19309355d0d21b6a64. Text cache verification passed with 393 exact prompts. The complete text index hash is recorded in the sidecar manifest. State normalization was computed from execution states in training episodes only.
The H100 SigLIP short-batch parity fix repeats a batch smaller than four before encoding and discards padding outputs. Live/offline parity across one sample from every task and both views passed with maximum RMSE 0 and maximum absolute error 0; the parity report is feature_parity.json.
The corrected text index hash is 627942ee144e1ea520a3ffaabee4e880231f63f610f17c394d65d950cf9f6a9e; pooled text hash is da4a754fb1878f91f76e5ac0001ada8f2852c524adb7a72f746af936aec1e22e; wrist manifest hash is 60ba47279af9ef70ac4100fc3b2170b34f6cfb9db5aea4107191a700c02dd2a4; state manifest hash is 6b01664c8717dd004f3e20eb19ac376c34e4531112d2daea8009f1e0be753f0c.
5. Timing-shortcut controls
The corrected run removes the dominant modulo-16 label leak: all trainable anchors share the deployment grid. The automated controls are in timing_baselines.json.
| Control on offline-dev | AP | AUROC |
|---|---|---|
| Class-prevalence/all-tied baseline | 0.278689 | 0.500000 |
| Anchor observation modulo 16 | 0.284294 | 0.506818 |
| Active-duration/controller chunks | 0.631068 | 0.841756 |
| Joint timing-only | 0.708675 | 0.908315 |
| Joint timing + task/phase | 0.864140 | 0.967499 |
The mandatory modulo-16 stop check passes: modulo-only AUROC is close to chance and below the 0.55 stop threshold. Duration remains a legitimate predictive cue and is reported separately; therefore the full model is not described as pure execution evidence based on AP alone.
6. Model, objective, and batching fixes
CommitReadinessHead uses frozen visual/text/state inputs, role/view/spatial/relative-age embeddings, one state-transformer layer, and two residual visual cross-attention blocks. The primary loss is BCE with ranking_loss_weight=0.0. Ranking remains available only for a later ablation.
The corrected trainer:
- uses one H100, per-device batch 64, accumulation 1, effective global batch exactly 64;
- drops incomplete batches (
drop_last=true), yielding 197 updates/epoch from 12,613 train rows (12,608 consumed per epoch); - derives warmup and cosine scheduler length from actual loader length: 197 warmup steps and 3,940 total updates;
- checks finite losses/gradients and records trainable parameter count;
- saves
last.pt,best_auprc.pt,best_nll.pt,best_pairwise_accuracy.pt, metrics, config, manifest, and signature.
The metric implementation groups equal scores before ROC/PR integration. Tied scores are permutation-invariant; one-class subsets are handled explicitly. This fixes the previous order-dependent AUROC/AP bug.
The immutable experiment signature covers canonical resolved config, CommitWAM/shared implementation files, all data/cache manifests and hashes, feature-encoder revision/config, and model architecture. It is stored in run_manifest.json, experiment_signature.json, and every checkpoint. Resume validates the signature before writing the output directory. The corrected production signature is:
480a8c8b16dedbdf7d7ba7ebd6dd8c7cf244b5fb64250ad2db58283b245f393e
7. Tests and preflight
The following checks passed before production training:
- Python compilation/import and shell syntax checks;
- Ruff format/check;
- 32 CommitWAM unit tests, including metric ties and permutation invariance;
- split-membership and target-corruption fail-closed tests;
- grid-alignment, causality, active-start, and feature allow-list tests;
- cache/hash and front/wrist/state/text parity checks;
- configuration invariants and BCE-only scope checks;
- checkpoint save/load and signature mismatch/resume no-overwrite tests;
- drop-last/effective-batch tests;
- corrected sidecar full audit and timing diagnostics.
The old four-H100 smoke artifacts are not used as evidence for this run because their implementation hash predates the corrected trainer. A final-code one-H100 smoke was run separately before production.
8. Final-code GPU smoke
Smoke output: .../grid_bce_seed7_1gpu_smoke_v2 (separate from production).
- one H100, batch 64, accumulation 1, effective global batch 64;
- 20 optimizer updates and a complete offline-dev evaluation;
- finite loss/gradients, no OOM/NaN;
- peak allocated 709,564,928 bytes (
0.661 GiB), peak reserved 918,552,576 bytes (0.855 GiB); - runtime 7.194 s, data wait 1.163 s, evaluation 3.004 s;
- AP 0.5272665, AUROC 0.8159293, NLL 0.6495998;
- checkpoint save/load passed;
- resumed two more updates with the same signature and unchanged manifest hash.
Production started from fresh random CommitReadinessHead initialization, not from this smoke or job 811950.
9. Production training result
Allocation and execution:
- Slurm allocation
815819, nodeevc104, one H100 80GB HBM3 (CUDA_VISIBLE_DEVICES=0); - 16 CPUs and approximately 1 TiB RAM; 5-hour interactive allocation;
- no
sbatchcall and no second allocation; - runtime feature transport used exact-manifest copies on node-local
/tmp/commitwam_815819.
Command:
source scripts/newton/robomme_env.sh
CUDA_VISIBLE_DEVICES=0 python -u -m fastwam.commitwam.train \
--config configs/commitwam/v01_grid_bce.yaml
Output:
/groups/yshang/an221229/checkpoints/FastWAM_Memory/runs/commitwam_v01/grid_bce_seed7_1gpu
The run completed all 20 epochs and exactly 3,940 optimizer updates. It used 12,613 train rows, 1,464 dev rows, batch 64, accumulation 1, drop-last, 197 updates/epoch, 197 warmup steps, BCE only, and 4,416,513 trainable parameters. Runtime was 350.23 s; peak allocated/reserved GPU memory was 709,564,928/920,649,728 bytes. No incomplete optimizer update, NaN/Inf, CUDA, or NCCL error occurred.
The best checkpoint was selected with corrected tied-score metrics:
best_auprc.pt: AP 0.9978101909060466, AUROC 0.9991748830213905, NLL 0.042958951610685595 (epoch 9);- final epoch 20: AP 0.9967723547290855, AUROC 0.9988058433600714, NLL 0.07935184241326779;
best_nll.ptandlast.ptare present.
10. Post-training feature diagnostics
The complete result is best_auprc_feature_ablations.json, evaluated on all 1,464 offline-dev rows:
| Input condition | AP | AUROC | NLL |
|---|---|---|---|
| Normal | 0.997810 | 0.999175 | 0.042959 |
| Relative-age zeroed | 0.947277 | 0.985501 | 0.215067 |
| Visual values zeroed | 0.990384 | 0.996668 | 0.089159 |
| State values zeroed | 0.672649 | 0.892276 | 1.146154 |
| Text embedding zeroed | 0.988629 | 0.995601 | 0.135911 |
The normal model beats the joint timing+task/phase control (0.864 AP), while the ablations show that state/history evidence is essential and timing is contributory rather than a modulo-grid-only explanation. This remains an offline proxy-label result, not a claim of physical completion or closed-loop success.
The old job-811950 read-only diagnostics are preserved separately. On its original-style evaluation, normal AP was 0.999717, age-zero 0.992485, visual-zero 0.998913, state-zero 0.753653, text-zero 0.997096, and controller-grid-only AP 0.993777. These numbers are not directly comparable to the corrected grid sidecar and are included only to document the clock-confounded run.
11. Reproducibility and explicit confirmations
- Training code commit:
4f6e8494dcd6bd576e8f894be5e760ff0aa18f47(Add verified node-local feature transport); final report changes may be committed separately. - Immutable implementation component hash in the production signature:
5a495974bb50d4ce62c1f1e196aba328d9c314eacdeeafde71e8a554a21885d3. - Branch:
CommitWAM-v01. - The training run recorded a dirty worktree because the user-owned untracked
paper/directory existed; it was not touched. - Production experiment signature:
480a8c8b16dedbdf7d7ba7ebd6dd8c7cf244b5fb64250ad2db58283b245f393e. - Job 811950 was untouched.
- Corrected training started from scratch.
- The current interactive allocation was used; no
sbatchjob was submitted. - Qwen, FastWAM, Memory Expert, SigLIP, and text encoder were frozen.
- Only
CommitReadinessHeadwas trained. - No human labels, candidate text, WAIT/HOLD, asynchronous planner, or test-set tuning was used.
- No off-grid BCE positives were used.
- No undersized optimizer update was used; effective global batch was exactly 64.
Closed-loop FIFO integration, validation-only commit-threshold selection, and official RoboMME test evaluation remain the next evaluation-server step. No official test result is claimed by this training report.