Update README.md
Browse files
README.md
CHANGED
|
@@ -13,7 +13,7 @@ tags:
|
|
| 13 |
- wan2.1
|
| 14 |
---
|
| 15 |
|
| 16 |
-
# LiveWan β streaming text
|
| 17 |
|
| 18 |
> **Unofficial community project.** Not affiliated with, endorsed by, or produced
|
| 19 |
> by Alibaba Group or the Wan-Video team. Built on their Apache-2.0
|
|
@@ -49,7 +49,7 @@ sentence.
|
|
| 49 |
| `out/world_p{0,44,60,82}.pt` | 4 Γ 59 MB | the four cached evaluation worlds β skip base-model generation entirely |
|
| 50 |
| `samples/` | 16 MB | reference clips and analysis filmstrips from those worlds |
|
| 51 |
|
| 52 |
-
**To continue the run
|
| 53 |
download these unless you intend to train.
|
| 54 |
|
| 55 |
| path | size | what |
|
|
@@ -90,44 +90,14 @@ browser stream.
|
|
| 90 |
**`--prompt-idx` must match the world**: `world_p60.pt` goes with `--prompt-idx 60`.
|
| 91 |
Indices run 0β95; anything above fails.
|
| 92 |
|
| 93 |
-
## Why step 2250 and not the final step 3000
|
| 94 |
-
|
| 95 |
-
The run went to 3000 iterations. Step 3000 is **not** the released checkpoint: it
|
| 96 |
-
more than doubles the sharpness-decay deviation against 2250, almost entirely on
|
| 97 |
-
world 60, where sharpness *rises* 30% across a 30-second stream.
|
| 98 |
-
|
| 99 |
-
| arm | step | mean \|decayβ1\| | mean \|drift\| | mean motion | min real-time |
|
| 100 |
-
|---|---|---|---|---|---|
|
| 101 |
-
| base | 2000 | 0.0775 | 0.0160 | 1.05 | 2.72Γ |
|
| 102 |
-
| prior run | 400 | 0.1325 | 0.0555 | 2.69 | 2.75Γ |
|
| 103 |
-
| this run | 750 | 0.0750 | 0.0915 | 1.91 | 2.49Γ |
|
| 104 |
-
| this run | 1500 | 0.0725 | 0.0580 | 1.93 | 2.43Γ |
|
| 105 |
-
| **this run** | **2250** | **0.0575** | **0.0178** | 2.17 | 2.77Γ |
|
| 106 |
-
| this run | 3000 | 0.1275 | 0.0240 | 2.21 | 2.76Γ |
|
| 107 |
-
|
| 108 |
-
(Real-time column as originally published, against a 25 fps budget; against the
|
| 109 |
-
native 16 fps it is 3.8β4.4Γ.)
|
| 110 |
-
|
| 111 |
-
Checkpoint quality here is **not monotone** β an earlier run went β0.024 β β0.256 β
|
| 112 |
-
β0.574 β β0.115 on drift across four consecutive checkpoints. The honest
|
| 113 |
-
counter-argument for step 3000, which is sharper and has more motion, is in the
|
| 114 |
-
code repo's `docs/HANDOFF.md` Β§1.
|
| 115 |
-
|
| 116 |
-
Step 3000 ships here as `checkpoints/t14b_b64/latest.pt` so the comparison is
|
| 117 |
-
yours to make rather than to take on trust. Swap it into the command above and
|
| 118 |
-
watch world 60 β that is where the two checkpoints actually differ.
|
| 119 |
-
|
| 120 |
-
`_noema` means the EMA copy was stripped to halve the file, 10.6 β 5.3 GB. This
|
| 121 |
-
changes nothing for inference: every evaluation in this project ran with
|
| 122 |
-
`use_ema: false`.
|
| 123 |
|
| 124 |
## Training
|
| 125 |
|
| 126 |
3000 iterations of SF-DMD distillation from a Wan2.1-T2V-14B teacher into a
|
| 127 |
Wan2.1-T2V-1.3B student, effective batch 64 (8ΓH200, accum 8, FSDP), **41.6 hours**
|
| 128 |
at 66.5 s/it, zero interventions. Losses do not decrease in this trainer and should
|
| 129 |
-
not
|
| 130 |
-
strengthening opponent.
|
| 131 |
|
| 132 |
## Continuing the run
|
| 133 |
|
|
|
|
| 13 |
- wan2.1
|
| 14 |
---
|
| 15 |
|
| 16 |
+
# LiveWan β streaming text 2 video, 3000 steps
|
| 17 |
|
| 18 |
> **Unofficial community project.** Not affiliated with, endorsed by, or produced
|
| 19 |
> by Alibaba Group or the Wan-Video team. Built on their Apache-2.0
|
|
|
|
| 49 |
| `out/world_p{0,44,60,82}.pt` | 4 Γ 59 MB | the four cached evaluation worlds β skip base-model generation entirely |
|
| 50 |
| `samples/` | 16 MB | reference clips and analysis filmstrips from those worlds |
|
| 51 |
|
| 52 |
+
**To continue the run** Not needed for inference. do not
|
| 53 |
download these unless you intend to train.
|
| 54 |
|
| 55 |
| path | size | what |
|
|
|
|
| 90 |
**`--prompt-idx` must match the world**: `world_p60.pt` goes with `--prompt-idx 60`.
|
| 91 |
Indices run 0β95; anything above fails.
|
| 92 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 93 |
|
| 94 |
## Training
|
| 95 |
|
| 96 |
3000 iterations of SF-DMD distillation from a Wan2.1-T2V-14B teacher into a
|
| 97 |
Wan2.1-T2V-1.3B student, effective batch 64 (8ΓH200, accum 8, FSDP), **41.6 hours**
|
| 98 |
at 66.5 s/it, zero interventions. Losses do not decrease in this trainer and should
|
| 99 |
+
not (the critic is retrained every step, so the generator holds position against a
|
| 100 |
+
strengthening opponent).
|
| 101 |
|
| 102 |
## Continuing the run
|
| 103 |
|