File size: 12,105 Bytes
3b2f112 c07a08a 3b2f112 c07a08a 3b2f112 c07a08a 3b2f112 c07a08a 3b2f112 a392c86 c07a08a 3b2f112 a392c86 3b2f112 a392c86 3b2f112 a392c86 3b2f112 a392c86 3b2f112 c07a08a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 | # Training code — CoMind 2-view generation (single-ego setup)
This bundle is our current 2-view egocentric video training stack (Cosmos-Predict2.5, `predict2_multiview`).
It is set up for a **2-view actor-actor** model with a lot of cross-view conditioning. You are going to
**strip it down to a single-ego generation setup that still outputs 2 views**, trained **from scratch on the
CoMind dataset**. This README lists exactly what to change and what to watch out for.
--------------------------------------------------------------------------------
## 0. What the current model conditions on (the parts you will REMOVE / keep)
Per view, each generated clip is conditioned on:
| Signal | Key | What it is | Single-ego plan |
|---|---|---|---|
| warped RGB | `control_input_warped` | RGB warped into the target view from a **combined pool (both views' past)** | **KEEP the channel, but change the DATA**: warp only from **that view's own past context** |
| pose | `control_input_pose` | skeletons. `person_pose=True` = identity-colored, **both people** visible | **KEEP the channel, change the DATA**: render **only the wearer's own** pose |
| in-context references | `reference_frames` (+ `reference_pose`, `reference_cam_w2c`) | clean appearance/pose anchor frames appended as extra tokens | **REMOVE entirely** |
| camera Plücker | `plucker_map` | per-pixel rays. Currently in a **shared cross-view canonical frame (view0 frame0)** | **KEEP** (`enable_plucker=True`), but change the canonical to **each view's OWN first frame** — see §1f |
| reference Plücker | `reference_plucker_map` | posed refs' rays | **REMOVE** (`enable_reference_plucker=False`; no refs) |
| view embedding | net `view_embeddings` | per-view identity added to tokens | **REMOVE** (drop `concat_view_embedding`) |
| cross-view self-attention | token layout `B (V·t) …` | both views attend each other | **Leave as-is** — you do NOT need to separate it; unified attention is fine, and with all cross-view conditioning removed the two views are effectively independent anyway |
So the single-ego model keeps: **warped_cond (own past) + own pose + camera Plücker (own-first-frame canonical)**,
generated per view, no refs / no reference-Plücker / no view-embedding, trained from the Cosmos base 2B.
--------------------------------------------------------------------------------
## 1. Code changes (all in `cosmos_predict2/_src/predict2_multiview/`)
Make a new experiment function in
`configs/vid2vid/experiment/nymeria_pose_2actor.py` (copy `comind_actoractor_personpose_shared` as a start),
and flip these toggles. The flags exist already — you're just turning conditioning OFF.
### 1a. model config (`models/multiview_pose_model_rectified_flow.py` fields, set in the experiment)
```python
cfg["model"]["config"]["num_reference_frames"] = 0 # was 4
cfg["model"]["config"]["enable_reference_plucker"] = False # was True (no refs)
cfg["model"]["config"]["enable_reference_pose"] = False # was True (no refs)
cfg["model"]["config"]["enable_plucker"] = True # KEEP camera Plücker (own-first-frame canonical, §1f)
```
### 1b. net config (`networks/multiview_pose_dit.py`, set via `cfg["model"]["config"]["net"].update(...)`)
```python
cfg["model"]["config"]["net"].update(
enable_reference_frames = False,
num_reference_frames = 0,
shared_reference = False,
enable_reference_plucker= False,
enable_reference_pose = False,
enable_plucker = True, # KEEP Plücker embedder (6-ch ray map -> zero-init add)
concat_view_embedding = False, # <-- drops the view embedding
)
```
`pose_mode` stays `"vae_concat"` (pose is VAE-encoded and concatenated into `cond_embedder`; keep it). The
`cond_embedder` in-channels auto-compute as `warped_latent(16) + visibility(1) + pose_latent(16)` — leave that.
### 1f. Plücker canonical frame = each view's OWN first frame — **already baked in the data, NO code change**
We want each view's Plücker rays expressed relative to its OWN first frame (not a shared cross-view frame).
**The `.single.npz` data already bakes this**: `L_w2c[0]` and `H_w2c[0]` are the identity matrix
(`cam_frame: per_view_canonical_frame0`), i.e. every view's camera poses are already canonicalized to its own
frame0. The model's existing `_preprocess` (~line 185) computes `c2w_ref = inv(w2c[:, 0, 0])`; since view0
frame0 is identity, `c2w_ref = I` and the shared transform is a **no-op**, so each view's rays come out in its
own already-canonical frame — exactly what we want. **Therefore: keep the model Plücker code AS-IS. Do NOT add
a per-view-canonical loop — the data already did it (adding one would be redundant).** Just keep
`enable_plucker=True`, `enable_reference_plucker=False`, and drop the `reference_plucker_map` branch (refs off).
### 1c. checkpoint = from scratch
```python
cfg["checkpoint"]["load_path"] = _base_2b_multiview_ckpt() # Cosmos base 2B, NOT a warm-start
cfg["checkpoint"]["strict_resume"] = False
```
Because refs/Plücker/view-emb are gone and the pose/warp embedders are zero-init or freshly-shaped, **iter-1
loss will start HIGH** (~base video prior) — that's correct for a from-scratch run, unlike a warm-start.
### 1d. data loader
Point `override /data_train` and `override /data_val` at a CoMind data config that:
- reads **your regenerated** CoMind clips (own-past warp + own-only pose),
- sets `num_reference_frames=0`, `shared_reference=False`, `emit_reference_pose=False` (do NOT emit any
`reference_*`), `person_pose=` your choice (you're rendering own-only pose, so the identity-color flag is moot).
The CoMind loader lives in `datasets/comind_pairs.py` / the actor-actor path in `datasets/nymeria_pairs.py`
(`get_nymeria_actor_actor_loader`, class `NymeriaActorActorDataset`). Register a new `cs.store(...)` data
config with the flags above.
### 1e. validation viz
`callbacks/nymeria_validation_viz.py`: set `num_reference_frames=0, shared_reference=False,
emit_reference_pose=False` in the viz kwargs too, or the viz dataset build will look for refs that aren't there.
--------------------------------------------------------------------------------
## 2. DATA — the ablation clips are the sibling `.single.npz` (already being generated)
The single-ego ablation data is stored as a **sibling `<clip>.single.npz`** next to the untouched originals
(`<clip>.npz`, `.campose.npz`, `.refs.npz`). The `.single.npz` is **self-contained** for training and its keys
match the loader's existing convention. Contents (JPEG byte-arrays for image streams, same as the originals):
```
L_/H_target GT video (JPEG) -> video
L_/H_warped OWN-only warp (srcview==0 kept; ==1 -> black hole) -> control_input_warped
L_/H_pose SELF-only skeleton (own hands only) -> control_input_pose
L_/H_vis_packed + vis_shape packed visibility (hole frames vis=False) -> control_input_visibility
L_/H_w2c (77,4,4) PER-VIEW canonical (each view frame0 = identity) -> camera_w2c
K_leader/K_helper (3,3) intrinsics -> camera_K
L_/H_srcf source frame idx (hole=-1) [loader ignores]
warp_mode/pose_mode/cam_frame/n_hole_* markers [loader ignores]
```
**IMPORTANT — source strictly from `.single.npz`.** The original `clip.npz` ALSO has `L_warped/L_pose/
L_vis_packed`, but those are the OLD combined-pool warp + both-people pose. Read warped/pose/vis (and w2c/K)
**only from `.single.npz`**. Since `.single.npz` also carries `L_target`, the simplest wiring is to use
`.single.npz` as the sole source and not open `clip.npz` at all (keypoints aren't read by training). No
`.refs.npz` (references are off).
### 2a. Loader adaptation (small — the format already matches)
Start from `datasets/comind_pairs.py` (`ComindActorActorDataset`), which already reads L_/H_ keys as JPEG
streams. Change:
- **Point the npz path at `<clip>.single.npz`** (and use it for the camera reads too — it has `L_w2c`/
`K_leader`). Drop the `.refs.npz` requirement in the eligibility loop and don't open refs.
- **`person_pose=False`** — `.single.npz` has `L_pose` (self-only), NOT `L_pose_person`; with
`person_pose=True` the loader would KeyError.
- **`num_reference_frames=0`** (R=0) so the reference block is skipped.
- Everything else is unchanged: `_stack_jpeg` decode, `np.unpackbits(vis_packed)` with `vis_shape=(77,504,504)`,
`_fit_poses(w2c,(77))`, `_resize` to your train resolution — all already match the `.single.npz` layout.
Verified against a real file (`62531cc1_006159.single.npz`, 28.5MB): warped/pose/target are JPEG object arrays;
`vis_shape=(77,504,504)` with packed-bit capacity exactly `77*504*504`; `L_w2c[0]==H_w2c[0]==identity`
(per-view canonical). So with the three config flips above, the existing loader/decoder path lines up 1:1.
--------------------------------------------------------------------------------
## 3. Things to know / gotchas
- **Env / launch**: `sh/train_nymeria_longer.sh`, run as
`CUDA_VISIBLE_DEVICES=0,1,2,3 EXP=<your_exp_name> NPROC=4 bash sh/train_nymeria_longer.sh`.
Conda env `ego_dh`. Needs `COSMOS_QWEN_TOKENIZER_DIR=/data/cosmos_reason1_7b` for the online text encoder,
`IMAGINAIRE_OUTPUT_ROOT=<out>`.
- **Online text encoding**: CoMind emits `ai_caption`; the model text-encodes it online each step (Qwen). Keep
the tokenizer dir env set or conditioning will crash on `t5_text_embeddings=None`.
- **DCP checkpoints** are reshardable across NPROC (4↔3↔2 resume works). `save_iter` in the experiment.
- **register the experiment**: append your function to the `experiments = [...]` list at the bottom of
`nymeria_pose_2actor.py`, and confirm it builds:
`python -c "..."` overriding `experiment=<name>` (see how the file builds configs) before launching 4-GPU.
- **`state_t`**: latent temporal length = `1 + (T-1)//4`. For 77 frames it's 20. It appears in the model config;
keep it consistent between data (`num_video_frames`) and model.
- **AR / noise recipe**: `predict2/models/video2world_model_rectified_flow.py` has an optional
`noisy_conditioning_*` + `conditional_frames_probs` recipe (teacher-forcing for autoregressive rollout). It is
OFF by default (`{1:1.0}`, prob 0). Ignore it unless you want AR — not needed for the single-ego baseline.
- **Cross-attention**: you asked whether to separate it — you don't need to. The DiT already runs unified
self-attention over `B (V·t) H W D`; with no cross-view conditioning the two views just don't share useful
signal, which is exactly the single-ego behavior. Leaving it unified is simplest and correct.
--------------------------------------------------------------------------------
## 4. File map (what's in this bundle)
```
cosmos_predict2/_src/predict2_multiview/
configs/vid2vid/experiment/nymeria_pose_2actor.py # all experiments (copy comind_* -> your single-ego exp)
configs/vid2vid/defaults/{conditioner,data,model}.py
datasets/nymeria_pairs.py # NymeriaActorActorDataset + loaders + cs.store data configs
datasets/comind_pairs.py # CoMind dataset/loader
networks/multiview_pose_dit.py # the DiT: all the enable_* / view-emb / reference embedders
models/multiview_pose_model_rectified_flow.py # conditioning preprocessing (refs/plucker/refpose/depth)
models/multiview_vid2vid_model_rectified_flow.py
callbacks/nymeria_validation_viz.py # validation grid
predict2/models/video2world_model_rectified_flow.py # base rectified-flow denoise (+optional AR noise)
sh/train_nymeria_longer.sh # launch
```
Summary of the diff you need: **turn OFF** `enable_reference_frames`, `enable_reference_plucker`,
`enable_reference_pose`, `num_reference_frames=0`, `shared_reference=False`, `concat_view_embedding=False`;
**KEEP `enable_plucker=True` but switch its canonical to each view's own first frame (§1f)**; **from base 2B**;
and feed **CoMind clips rebuilt with own-past warp + own-only pose**. Keep warped_cond + pose + Plücker
channels and unified cross-view attention.
|