File size: 12,105 Bytes
3b2f112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c07a08a
 
3b2f112
 
 
c07a08a
 
3b2f112
 
 
 
 
 
 
 
 
 
 
c07a08a
 
 
3b2f112
 
 
 
 
 
 
 
 
 
c07a08a
3b2f112
 
 
 
 
 
a392c86
 
 
 
 
 
 
 
 
c07a08a
3b2f112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a392c86
3b2f112
a392c86
 
 
3b2f112
a392c86
 
 
 
 
 
 
 
 
 
3b2f112
a392c86
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3b2f112
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c07a08a
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
# Training code — CoMind 2-view generation (single-ego setup)

This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, `predict2_multiview`).
It is set up for a **2-view actor-actor** model with a lot of cross-view conditioning. You are going to
**strip it down to a single-ego generation setup that still outputs 2 views**, trained **from scratch on the
CoMind dataset**. This README lists exactly what to change and what to watch out for.

--------------------------------------------------------------------------------
## 0. What the current model conditions on (the parts you will REMOVE / keep)

Per view, each generated clip is conditioned on:
| Signal | Key | What it is | Single-ego plan |
|---|---|---|---|
| warped RGB | `control_input_warped` | RGB warped into the target view from a **combined pool (both views' past)** | **KEEP the channel, but change the DATA**: warp only from **that view's own past context** |
| pose | `control_input_pose` | skeletons. `person_pose=True` = identity-colored, **both people** visible | **KEEP the channel, change the DATA**: render **only the wearer's own** pose |
| in-context references | `reference_frames` (+ `reference_pose`, `reference_cam_w2c`) | clean appearance/pose anchor frames appended as extra tokens | **REMOVE entirely** |
| camera Plücker | `plucker_map` | per-pixel rays. Currently in a **shared cross-view canonical frame (view0 frame0)** | **KEEP** (`enable_plucker=True`), but change the canonical to **each view's OWN first frame** — see §1f |
| reference Plücker | `reference_plucker_map` | posed refs' rays | **REMOVE** (`enable_reference_plucker=False`; no refs) |
| view embedding | net `view_embeddings` | per-view identity added to tokens | **REMOVE** (drop `concat_view_embedding`) |
| cross-view self-attention | token layout `B (V·t) …` | both views attend each other | **Leave as-is** — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway |

So the single-ego model keeps: **warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical)**,
generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B.

--------------------------------------------------------------------------------
## 1. Code changes (all in `cosmos_predict2/_src/predict2_multiview/`)

Make a new experiment function in
`configs/vid2vid/experiment/nymeria_pose_2actor.py` (copy `comind_actoractor_personpose_shared` as a start),
and flip these toggles. The flags exist already — you're just turning conditioning OFF.

### 1a. model config (`models/multiview_pose_model_rectified_flow.py` fields, set in the experiment)
```python
cfg["model"]["config"]["num_reference_frames"]      = 0       # was 4
cfg["model"]["config"]["enable_reference_plucker"]  = False   # was True (no refs)
cfg["model"]["config"]["enable_reference_pose"]     = False   # was True (no refs)
cfg["model"]["config"]["enable_plucker"]            = True    # KEEP camera Plücker (own-first-frame canonical, §1f)
```

### 1b. net config (`networks/multiview_pose_dit.py`, set via `cfg["model"]["config"]["net"].update(...)`)
```python
cfg["model"]["config"]["net"].update(
    enable_reference_frames = False,
    num_reference_frames    = 0,
    shared_reference        = False,
    enable_reference_plucker= False,
    enable_reference_pose   = False,
    enable_plucker          = True,    # KEEP Plücker embedder (6-ch ray map -> zero-init add)
    concat_view_embedding   = False,   # <-- drops the view embedding
)
```
`pose_mode` stays `"vae_concat"` (pose is VAE-encoded and concatenated into `cond_embedder`; keep it). The
`cond_embedder` in-channels auto-compute as `warped_latent(16) + visibility(1) + pose_latent(16)` — leave that.

### 1f. Plücker canonical frame = each view's OWN first frame — **already baked in the data, NO code change**
We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame).
**The `.single.npz` data already bakes this**: `L_w2c[0]` and `H_w2c[0]` are the identity matrix
(`cam_frame: per_view_canonical_frame0`), i.e. every view's camera poses are already canonicalized to its own
frame0. The model's existing `_preprocess` (~line 185) computes `c2w_ref = inv(w2c[:, 0, 0])`; since view0
frame0 is identity, `c2w_ref = I` and the shared transform is a **no-op**, so each view's rays come out in its
own already-canonical frame — exactly what we want. **Therefore: keep the model Plücker code AS-IS. Do NOT add
a per-view-canonical loop — the data already did it (adding one would be redundant).** Just keep
`enable_plucker=True`, `enable_reference_plucker=False`, and drop the `reference_plucker_map` branch (refs off).

### 1c. checkpoint = from scratch
```python
cfg["checkpoint"]["load_path"]      = _base_2b_multiview_ckpt()   # Cosmos base 2B, NOT a warm-start
cfg["checkpoint"]["strict_resume"]  = False
```
Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, **iter-1
loss will start HIGH** (~base video prior) — that's correct for a from-scratch run, unlike a warm-start.

### 1d. data loader
Point `override /data_train` and `override /data_val` at a CoMind data config that:
- reads **your regenerated** CoMind clips (own-past warp + own-only pose),
- sets `num_reference_frames=0`, `shared_reference=False`, `emit_reference_pose=False` (do NOT emit any
  `reference_*`), `person_pose=` your choice (you're rendering own-only pose, so the identity-color flag is moot).

The CoMind loader lives in `datasets/comind_pairs.py` / the actor-actor path in `datasets/nymeria_pairs.py`
(`get_nymeria_actor_actor_loader`, class `NymeriaActorActorDataset`). Register a new `cs.store(...)` data
config with the flags above.

### 1e. validation viz
`callbacks/nymeria_validation_viz.py`: set `num_reference_frames=0, shared_reference=False,
emit_reference_pose=False` in the viz kwargs too, or the viz dataset build will look for refs that aren't there.

--------------------------------------------------------------------------------
## 2. DATA — the ablation clips are the sibling `.single.npz` (already being generated)

The single-ego ablation data is stored as a **sibling `<clip>.single.npz`** next to the untouched originals
(`<clip>.npz`, `.campose.npz`, `.refs.npz`). The `.single.npz` is **self-contained** for training and its keys
match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals):

```
L_/H_target        GT video (JPEG)                         -> video
L_/H_warped        OWN-only warp (srcview==0 kept; ==1 -> black hole)  -> control_input_warped
L_/H_pose          SELF-only skeleton (own hands only)      -> control_input_pose
L_/H_vis_packed + vis_shape   packed visibility (hole frames vis=False) -> control_input_visibility
L_/H_w2c (77,4,4)  PER-VIEW canonical (each view frame0 = identity)     -> camera_w2c
K_leader/K_helper (3,3)   intrinsics                        -> camera_K
L_/H_srcf          source frame idx (hole=-1)  [loader ignores]
warp_mode/pose_mode/cam_frame/n_hole_*   markers [loader ignores]
```

**IMPORTANT — source strictly from `.single.npz`.** The original `clip.npz` ALSO has `L_warped/L_pose/
L_vis_packed`, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K)
**only from `.single.npz`**. Since `.single.npz` also carries `L_target`, the simplest wiring is to use
`.single.npz` as the sole source and not open `clip.npz` at all (keypoints aren't read by training). No
`.refs.npz` (references are off).

### 2a. Loader adaptation (small — the format already matches)
Start from `datasets/comind_pairs.py` (`ComindActorActorDataset`), which already reads L_/H_ keys as JPEG
streams. Change:
- **Point the npz path at `<clip>.single.npz`** (and use it for the camera reads too — it has `L_w2c`/
  `K_leader`). Drop the `.refs.npz` requirement in the eligibility loop and don't open refs.
- **`person_pose=False`** — `.single.npz` has `L_pose` (self-only), NOT `L_pose_person`; with
  `person_pose=True` the loader would KeyError.
- **`num_reference_frames=0`** (R=0) so the reference block is skipped.
- Everything else is unchanged: `_stack_jpeg` decode, `np.unpackbits(vis_packed)` with `vis_shape=(77,504,504)`,
  `_fit_poses(w2c,(77))`, `_resize` to your train resolution — all already match the `.single.npz` layout.

Verified against a real file (`62531cc1_006159.single.npz`, 28.5MB): warped/pose/target are JPEG object arrays;
`vis_shape=(77,504,504)` with packed-bit capacity exactly `77*504*504`; `L_w2c[0]==H_w2c[0]==identity`
(per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1.

--------------------------------------------------------------------------------
## 3. Things to know / gotchas

- **Env / launch**: `sh/train_nymeria_longer.sh`, run as
  `CUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh`.
  Conda env `ego_dh`. Needs `COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b` for the online text encoder,
  `IMAGINAIRE_OUTPUT_ROOT=<out>`.
- **Online text encoding**: CoMind emits `ai_caption`; the model text-encodes it online each step (Qwen). Keep
  the tokenizer dir env set or conditioning will crash on `t5_text_embeddings=None`.
- **DCP checkpoints** are reshardable across NPROC (4↔3↔2 resume works). `save_iter` in the experiment.
- **register the experiment**: append your function to the `experiments = [...]` list at the bottom of
  `nymeria_pose_2actor.py`, and confirm it builds:
  `python -c "..."` overriding `experiment=<name>` (see how the file builds configs) before launching 4-GPU.
- **`state_t`**: latent temporal length = `1 + (T-1)//4`. For 77 frames it's 20. It appears in the model config;
  keep it consistent between data (`num_video_frames`) and model.
- **AR / noise recipe**: `predict2/models/video2world_model_rectified_flow.py` has an optional
  `noisy_conditioning_*` + `conditional_frames_probs` recipe (teacher-forcing for autoregressive rollout). It is
  OFF by default (`{1:1.0}`, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline.
- **Cross-attention**: you asked whether to separate it — you don't need to. The DiT already runs unified
  self-attention over `B (V·t) H W D`; with no cross-view conditioning the two views just don't share useful
  signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct.

--------------------------------------------------------------------------------
## 4. File map (what's in this bundle)

```
cosmos_predict2/_src/predict2_multiview/
  configs/vid2vid/experiment/nymeria_pose_2actor.py   # all experiments (copy comind_* -> your single-ego exp)
  configs/vid2vid/defaults/{conditioner,data,model}.py
  datasets/nymeria_pairs.py            # NymeriaActorActorDataset + loaders + cs.store data configs
  datasets/comind_pairs.py             # CoMind dataset/loader
  networks/multiview_pose_dit.py       # the DiT: all the enable_* / view-emb / reference embedders
  models/multiview_pose_model_rectified_flow.py   # conditioning preprocessing (refs/plucker/refpose/depth)
  models/multiview_vid2vid_model_rectified_flow.py
  callbacks/nymeria_validation_viz.py  # validation grid
predict2/models/video2world_model_rectified_flow.py  # base rectified-flow denoise (+optional AR noise)
sh/train_nymeria_longer.sh            # launch
```

Summary of the diff you need: **turn OFF** `enable_reference_frames`, `enable_reference_plucker`,
`enable_reference_pose`, `num_reference_frames=0`, `shared_reference=False`, `concat_view_embedding=False`;
**KEEP `enable_plucker=True` but switch its canonical to each view's own first frame (§1f)**; **from base 2B**;
and feed **CoMind clips rebuilt with own-past warp + own-only pose**. Keep warped_cond + pose + Plücker
channels and unified cross-view attention.