JonathanColetti commited on
Commit
8d2321d
Β·
verified Β·
1 Parent(s): 3c199e3

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +4 -34
README.md CHANGED
@@ -13,7 +13,7 @@ tags:
13
  - wan2.1
14
  ---
15
 
16
- # LiveWan β€” streaming text-to-video, step 2250
17
 
18
  > **Unofficial community project.** Not affiliated with, endorsed by, or produced
19
  > by Alibaba Group or the Wan-Video team. Built on their Apache-2.0
@@ -49,7 +49,7 @@ sentence.
49
  | `out/world_p{0,44,60,82}.pt` | 4 Γ— 59 MB | the four cached evaluation worlds β€” skip base-model generation entirely |
50
  | `samples/` | 16 MB | reference clips and analysis filmstrips from those worlds |
51
 
52
- **To continue the run β€” a further 28.3 GB.** Not needed for inference; do not
53
  download these unless you intend to train.
54
 
55
  | path | size | what |
@@ -90,44 +90,14 @@ browser stream.
90
  **`--prompt-idx` must match the world**: `world_p60.pt` goes with `--prompt-idx 60`.
91
  Indices run 0–95; anything above fails.
92
 
93
- ## Why step 2250 and not the final step 3000
94
-
95
- The run went to 3000 iterations. Step 3000 is **not** the released checkpoint: it
96
- more than doubles the sharpness-decay deviation against 2250, almost entirely on
97
- world 60, where sharpness *rises* 30% across a 30-second stream.
98
-
99
- | arm | step | mean \|decayβˆ’1\| | mean \|drift\| | mean motion | min real-time |
100
- |---|---|---|---|---|---|
101
- | base | 2000 | 0.0775 | 0.0160 | 1.05 | 2.72Γ— |
102
- | prior run | 400 | 0.1325 | 0.0555 | 2.69 | 2.75Γ— |
103
- | this run | 750 | 0.0750 | 0.0915 | 1.91 | 2.49Γ— |
104
- | this run | 1500 | 0.0725 | 0.0580 | 1.93 | 2.43Γ— |
105
- | **this run** | **2250** | **0.0575** | **0.0178** | 2.17 | 2.77Γ— |
106
- | this run | 3000 | 0.1275 | 0.0240 | 2.21 | 2.76Γ— |
107
-
108
- (Real-time column as originally published, against a 25 fps budget; against the
109
- native 16 fps it is 3.8–4.4Γ—.)
110
-
111
- Checkpoint quality here is **not monotone** β€” an earlier run went βˆ’0.024 β†’ βˆ’0.256 β†’
112
- βˆ’0.574 β†’ βˆ’0.115 on drift across four consecutive checkpoints. The honest
113
- counter-argument for step 3000, which is sharper and has more motion, is in the
114
- code repo's `docs/HANDOFF.md` Β§1.
115
-
116
- Step 3000 ships here as `checkpoints/t14b_b64/latest.pt` so the comparison is
117
- yours to make rather than to take on trust. Swap it into the command above and
118
- watch world 60 β€” that is where the two checkpoints actually differ.
119
-
120
- `_noema` means the EMA copy was stripped to halve the file, 10.6 β†’ 5.3 GB. This
121
- changes nothing for inference: every evaluation in this project ran with
122
- `use_ema: false`.
123
 
124
  ## Training
125
 
126
  3000 iterations of SF-DMD distillation from a Wan2.1-T2V-14B teacher into a
127
  Wan2.1-T2V-1.3B student, effective batch 64 (8Γ—H200, accum 8, FSDP), **41.6 hours**
128
  at 66.5 s/it, zero interventions. Losses do not decrease in this trainer and should
129
- not β€” the critic is retrained every step, so the generator holds position against a
130
- strengthening opponent.
131
 
132
  ## Continuing the run
133
 
 
13
  - wan2.1
14
  ---
15
 
16
+ # LiveWan β€” streaming text 2 video, 3000 steps
17
 
18
  > **Unofficial community project.** Not affiliated with, endorsed by, or produced
19
  > by Alibaba Group or the Wan-Video team. Built on their Apache-2.0
 
49
  | `out/world_p{0,44,60,82}.pt` | 4 Γ— 59 MB | the four cached evaluation worlds β€” skip base-model generation entirely |
50
  | `samples/` | 16 MB | reference clips and analysis filmstrips from those worlds |
51
 
52
+ **To continue the run** Not needed for inference. do not
53
  download these unless you intend to train.
54
 
55
  | path | size | what |
 
90
  **`--prompt-idx` must match the world**: `world_p60.pt` goes with `--prompt-idx 60`.
91
  Indices run 0–95; anything above fails.
92
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
93
 
94
  ## Training
95
 
96
  3000 iterations of SF-DMD distillation from a Wan2.1-T2V-14B teacher into a
97
  Wan2.1-T2V-1.3B student, effective batch 64 (8Γ—H200, accum 8, FSDP), **41.6 hours**
98
  at 66.5 s/it, zero interventions. Losses do not decrease in this trainer and should
99
+ not (the critic is retrained every step, so the generator holds position against a
100
+ strengthening opponent).
101
 
102
  ## Continuing the run
103