|
Download README.md from Cccccz/comparison: direct link, hf CLI and curl.
- Browser
- Download file 22.6 kB
-
https://huggingface.co/Cccccz/comparison/resolve/main/README.md
- Command line
-
hf download hf://Cccccz/comparison/README.md
-
curl -L -o README.md https://huggingface.co/Cccccz/comparison/resolve/main/README.md
22.6 kB
| # Cache-method comparison on 4-step block-causal video DiTs | |
| Four training-free caching methods (TeaCache, TaylorSeer, FlowCache, MotionCache) | |
| implemented on two 4-step autoregressive video base models (Self-Forcing and | |
| Causal-Forcing), swept to matched denoise-DiT speedups. | |
| ## Layout | |
| ``` | |
| repos/ the six upstream clones (reference implementations + base models) | |
| Self-Forcing/ base model 1 (guandeh17/Self-Forcing) | |
| Causal-Forcing/ base model 2 (thu-ml/Causal-Forcing) | |
| TeaCache/ ali-vilab/TeaCache | |
| TaylorSeer/ Shenyi-Z/TaylorSeer | |
| FlowCache/ mikeallen39/FlowCache (ICLR 2026) | |
| MotionCache/ MAC-AutoML/MotionCache (ICML 2026) | |
| cachelib/ the ports -- one implementation, both base models | |
| methods.py the four cache methods + baseline + calibration probe | |
| selective.py token-subset forward for the KV-cached causal DiT | |
| patch.py routes CausalWanModel._forward_inference through a method | |
| runner.py chunked denoising loop with cache hooks and denoise-only timing | |
| harness.py model loading, prompt loading, paired A/B measurement | |
| calibrate.py fits the TeaCache rescale polynomial for a base model | |
| sweep.py finds the parameter that hits a target speedup | |
| finalize.py re-targets on the timed prompts + held-out check (authoritative) | |
| overhead_bench.py per-step cost of a cached vs a full step (contention-robust) | |
| retime.py plain stopwatch re-measurement of the settled points | |
| verify.py re-measures operating points on held-out prompts | |
| run_compare.py single-configuration CLI | |
| summarize.py collects everything into one table | |
| results/ JSON output, plus SUMMARY.md / SUMMARY.csv | |
| videos/ generated mp4s per operating point, plus *_baseline references | |
| ``` | |
| Pipeline order: `calibrate.py` -> `sweep.py` -> `finalize.py` -> `overhead_bench.py` | |
| -> `retime.py` -> `summarize.py`. | |
| Weights are symlinked from the pre-existing checkouts rather than re-downloaded: | |
| | base | checkpoint | source | | |
| |---|---|---| | |
| | self_forcing | `checkpoints/self_forcing_dmd.pt` (`generator_ema`) | `../Self-Forcing` | | |
| | causal_forcing | `checkpoints/chunkwise/causal_forcing.pt` (`generator`) | `../Causal-Forcing` | | |
| Both also symlink `wan_models/` (Wan2.1-T2V-1.3B backbone, VAE, UMT5 text encoder). | |
| ## The setting | |
| Both base models are the same architecture: a block-causal Wan2.1-1.3B DiT that | |
| generates a video chunk at a time. With `num_frame_per_block=3` and 21 latent | |
| frames, a video is **7 chunks x 4 denoising steps = 28 DiT forwards**, plus one | |
| untimed KV-cache refresh pass per chunk (`context_noise`) that rewrites the | |
| chunk's KV entries from the clean latent. | |
| The experiment fixes the schedule to **`F ? ? F`**: step 0 and step 3 always run | |
| the full DiT, steps 1 and 2 are the ones a cache method may skip. This is | |
| equivalent to the upstream `ret_steps=1` / `cutoff_steps=num_steps-1` guards. | |
| **Speedup is measured on denoise-DiT time only** β the sum of the 28 denoising | |
| forwards, timed with CUDA events. Text encoding, VAE decode and the per-chunk | |
| KV-refresh pass are excluded. That makes 2.0x the arithmetic ceiling (14 of 28 | |
| forwards), which is why the 2.0x target is exactly "skip both middle steps | |
| everywhere". | |
| ## How each method was ported | |
| All four hook the same place: the 30-block loop inside | |
| `CausalWanModel._forward_inference`. The DiT preamble and the head/unpatchify tail | |
| always run. | |
| | method | decision granularity | what a skipped step reuses | knob | | |
| |---|---|---|---| | |
| | TeaCache | whole forward | the 30-block residual `x_out - x_in` from the last computed step | `thresh` | | |
| | FlowCache | per frame group inside the chunk | the same residual, banked per group | `thresh`, `group_size` | | |
| | TaylorSeer | whole forward | a **Taylor forecast** of each block's self-attn / cross-attn / FFN output | `interval`, `max_order` | | |
| | MotionCache | per token, motion-weighted | cached rows for un-selected tokens; selected tokens are recomputed | `thresh`, `weight_norm` | | |
| Skipping the block stack leaves the chunk's KV-cache entries holding the previous | |
| computed step's keys/values. That is safe here: every later step of the chunk | |
| overwrites the same slots, and the per-chunk KV-refresh pass rewrites them from | |
| the clean latent before the next chunk reads them. | |
| ### Selective (partial-token) forward | |
| FlowCache and MotionCache recompute only part of the chunk. On a KV-cached causal | |
| model that means, per self-attention layer: compute q/k/v for the selected rows | |
| only, scatter their fresh k/v into the slots the chunk owns (leaving unselected | |
| slots holding the previous step's k/v), and attend the selected queries against | |
| the full cache. This mirrors MotionCache's `forward_selective`, moved from | |
| SkyReels' persistent buffers onto Self-Forcing's rolling `kv_cache`. | |
| Validated: with *all* tokens selected the selective path is **bit-exact** against | |
| the stock full forward (max abs latent difference 0.0). The same check passes for | |
| TaylorSeer's recording path at `interval=1`. | |
| ### Deviations from upstream, and why | |
| 1. **Indicator.** `TeaCache4Wan2.1` derives its indicator from `e0` (the timestep | |
| modulation) alone. That is fine for bidirectional Wan, where one forward covers | |
| the whole video at a single timestep, but it *degenerates* here: every chunk | |
| runs the same four timesteps, so an `e0` indicator is bit-identical across | |
| chunks and the threshold cannot adapt to content at all. The default is | |
| therefore TeaCache's original definition β the relative L1 distance of the | |
| **timestep-modulated noisy input**, the tensor block 0 feeds to its | |
| self-attention, which is what TeaCache uses on HunyuanVideo and FLUX. | |
| `--indicator e0` restores the literal Wan port. | |
| 2. **Rescale polynomial.** TeaCache's shipped 4th-degree coefficients were fitted | |
| for a different model, schedule and indicator, so `calibrate.py` reruns | |
| TeaCache's own fitting procedure here. The fit is **degree 1, not 4**: on this | |
| model the input distance only spans ~[1.19, 1.36], where a degree-4 fit is | |
| ill-conditioned (coefficients ~10^3 with alternating signs) and MotionCache, | |
| which evaluates the polynomial *per token*, would extrapolate it far outside | |
| the fitted domain. Degree 1 is monotone and safe. All degrees are recorded in | |
| the coefficient JSON. | |
| 3. **FlowCache granularity.** FlowCache's contribution is that the units covered | |
| by one forward denoise at different rates and need independent policies. In | |
| SkyReels-V2 those units are the chunks of the diffusion-forcing window; here a | |
| forward covers exactly one chunk, so the corresponding units are the frame | |
| groups inside it (`group_size=1`, i.e. 3 groups). Its KV-cache-compression | |
| component is not ported: Self-Forcing already bounds its KV cache by local | |
| attention, so there is no growing cache to compress. | |
| 4. **TaylorSeer fractional interval.** TaylorSeer's schedule is a uniform integer | |
| refresh interval, which on an `F ? ? F` schedule can only produce 1.0x, 1.33x | |
| and 2.0x. `interval` is accepted as a float and realised as a deterministic | |
| per-chunk alternation between the two bracketing integer intervals | |
| (Bresenham-style), giving the requested *average* interval. This keeps | |
| TaylorSeer's content-independent schedule rather than bolting a threshold onto | |
| it. | |
| 5. **Fused forecast.** TaylorSeer's cached step is ~90 bandwidth-bound elementwise | |
| ops per forward. Upstream puts `@torch.compile` on the equivalent | |
| `wan_attention_cache_forward`; the port does the same. Without it the forecast | |
| ate most of the saving (1.72x measured at a 2.0x FLOPs setting); with it, it | |
| does not. Set `CACHELIB_NO_COMPILE=1` to disable. | |
| ## Scope change (Sep 4) | |
| FlowCache is frozen: its finished operating points stay in the tables, nothing new | |
| is generated for it. The continuing study is **three methods** (TeaCache, | |
| TaylorSeer, MotionCache) x **three base models** (Self-Forcing, Causal-Forcing, | |
| HY-WorldPlay) x **three points**: `FxxF` at ~1.3x, `FxxF` at 1.7-1.8x, and `Fxxx` | |
| at 2.7-3.0x -- all targeted on **measured** wall-clock speedup, not on the compute | |
| fraction. The reason is MotionCache: matching compute at 2.0x under `FxxF` forces | |
| its per-token selection to pick *zero* tokens, which makes it bit-identical to | |
| TeaCache and reproduces nothing. Operating points are now chosen as the fastest | |
| setting that still keeps the token-level decision active (>=25% of cacheable steps | |
| partial, >=5% of tokens selected on those steps); the activity statistics are | |
| recorded per video in ``cache_diagnostics``. | |
| ## Schedules beyond `FxxF` | |
| Every entry point takes `--schedule` (`sweep.py`, `finalize.py`, `run_compare.py`; | |
| the evaluation reads it from the `final_*.json` rows). `F` marks a step that always | |
| runs the full DiT, `x` one the method may serve from cache: `FxxF` is the study | |
| above and the default; `FFxx` gives TaylorSeer two computed points before its first | |
| forecast so the first-order term is actually exercised; `Fxxx` lifts the cool-down | |
| guard and raises the arithmetic ceiling to 4x. Operating points for `FFxx` | |
| (TaylorSeer, 1.3/1.6/2.0x) and `Fxxx` (the other three, ~3x) were searched on | |
| 3 prompts (`results/sweep_*_{FFxx,Fxxx}.json`, collected in | |
| `results/final_<base>_newsched.json`) and evaluated on Extended-251. | |
| **All results, speed and quality, are collected in `Experiment.md`** | |
| (regenerate with `python experiment_md.py`). The headline of the second round: | |
| caching the *last* denoising step -- which every `Fxxx` point and every `FFxx` | |
| point at >= 1.6x does -- costs 15-25 VBench points on both base models, while | |
| `FxxF` at the same compute costs 1-2. On a 4-step distilled model the final step | |
| is not optional. | |
| ## Measurement protocol | |
| This is a **shared node** and it was busy throughout. Other tenants' jobs move | |
| absolute timings by 30%+ and drift them *within* a single sweep, so a baseline | |
| measured once at the start silently inflates every later configuration. Four | |
| things follow, and they are the difference between numbers worth reading and | |
| numbers that are noise: | |
| 1. **Search on compute, not on time.** The parameter search runs on the compute | |
| fraction (how many of the 28 DiT forwards were actually computed), which is | |
| exactly deterministic given the prompt set and immune to contention. Targets | |
| are aimed at the compute budget directly, which also puts all four methods on | |
| identical compute β the right basis for a later quality comparison. Correcting | |
| the target by a measured overhead ratio was tried and abandoned: that ratio is | |
| itself a timing measurement, one probe returned 1.110 (physically impossible), | |
| and it dragged a whole method's operating points below target. | |
| 2. **Paired A/B.** Baseline and method run back to back on the same prompt, so | |
| drift between them is seconds rather than minutes. | |
| 3. **Minimum, not median, as the headline.** Contention is one-sided β a | |
| neighbour can only make a run slower β so the fastest observed run is the best | |
| estimate of the uncontended time, and the headline speedup is | |
| `min(baseline) / min(method)` over the pairs. The paired median is reported | |
| alongside but is biased *upward*, because the baseline run is longer and so | |
| more exposed to being hit by a spike; that is why it sometimes exceeded a | |
| method's arithmetic ceiling. | |
| 4. **Compute fraction from the timed runs.** A threshold tuned on one prompt | |
| subset and timed on another silently disagrees whenever it sits near a decision | |
| boundary. `finalize.py` reports the compute fraction of the very runs it timed, | |
| and re-tunes on those prompts if it misses. | |
| `finalize.py` also re-checks every point on a disjoint prompt slice, so a | |
| threshold that only hit its target by sitting inside a narrow band of the | |
| indicator distribution shows up as held-out drift. `verify.py` does the same | |
| check standalone against a sweep file. | |
| ## Reproducing | |
| ```bash | |
| PY=/local/zoubin/cz/envs/self_forcing/bin/python | |
| # 1. fit the rescale polynomial (once per base model) | |
| CUDA_VISIBLE_DEVICES=0 $PY calibrate.py --base self_forcing \ | |
| --num-prompts 12 --out results/coeff_self_forcing.json | |
| # 2. sweep one method to the three target speedups | |
| CUDA_VISIBLE_DEVICES=0 $PY sweep.py --base self_forcing --method teacache \ | |
| --coefficients results/coeff_self_forcing.json \ | |
| --targets 1.3,1.6,2.0 --save-video-dir videos \ | |
| --out results/sweep_self_forcing_teacache.json | |
| # 3. check the settings transfer to unseen prompts | |
| CUDA_VISIBLE_DEVICES=0 $PY verify.py \ | |
| --sweep results/sweep_self_forcing_teacache.json --prompt-offset 64 \ | |
| --out results/verify_self_forcing_teacache.json | |
| # 4. collect everything | |
| $PY summarize.py | |
| ``` | |
| `run_compare.py` runs a single configuration if you just want one number: | |
| ```bash | |
| CUDA_VISIBLE_DEVICES=0 $PY run_compare.py --base causal_forcing \ | |
| --method taylorseer --interval 2.5 --num-prompts 8 | |
| ``` | |
| ## Environment note | |
| **Sep 3:** the machine was restored again -- `/mnt/local_nvme` became `/local`, | |
| both environments lost their executable bits a second time (1632 files restored), | |
| the `vbench_eval` venv's `pyvenv.cfg`/interpreter links pointed at the old path | |
| (now symlinks into `/local/.../self_forcing`), and the caches moved to | |
| `/local/zoubin/cz/.cache/`. Paths in `eval/*.sh`, `eval/aggregate.py` and the | |
| per-prompt records were updated accordingly. The node is shared: other tenants' | |
| launches have repeatedly killed GPU processes of ours, so long runs go through | |
| `eval/retime_until_clean.sh`-style loops that only use quiet GPUs and resume. | |
| The conda environment at `/mnt/local_nvme/zoubin/cz/envs/self_forcing` had lost | |
| every executable bit (a bad restore left all files `rw-rw-r--`), so nothing in it | |
| could run β including Triton's bundled `ptxas`, which broke `torch.compile`. The | |
| exec bits were restored on files that are ELF binaries or start with `#!` | |
| (1772 files). No package contents were changed. | |
| ## Results | |
| Full tables: `results/SUMMARY.md` (regenerate with `python summarize.py`), machine | |
| readable in `results/SUMMARY.csv` and the per-stage JSONs. | |
| **Headline speedup** is the denoise-DiT speedup implied by the measured compute | |
| fraction and the measured per-step cost model (see Measurement protocol). The | |
| `stopwatch` column is the plain paired whole-video ratio; it agrees within its | |
| (large) spread but is not the number to quote from this node. | |
| ### Operating points found | |
| | base | method | knob | 1.3x | 1.6x | 2.0x | | |
| |---|---|---|---|---|---| | |
| | Self-Forcing | TeaCache | `thresh` | 0.6348 β **1.30x** | 1.2507 β **1.48x** | 1.2917 β **1.96x** | | |
| | Self-Forcing | FlowCache | `thresh` | 0.6360 β **1.30x** | 1.2500 β **1.51x** | 1.2917 β **1.96x** | | |
| | Self-Forcing | TaylorSeer | `interval` | 2.083 β **1.31x** | 2.583 β **1.59x** | 2.917 β **1.89x** | | |
| | Self-Forcing | MotionCache | `thresh` | 0.7560 β **1.29x** | 1.2461 β **1.60x** | 2.5000 β **1.96x** | | |
| | Causal-Forcing | TeaCache | `thresh` | 0.6491 β **1.30x** | 1.2760 β **1.51x** | 1.2917 β **1.96x** | | |
| | Causal-Forcing | FlowCache | `thresh` | 0.6488 β **1.26x** | 1.2760 β **1.61x** | 1.2917 β **1.95x** | | |
| | Causal-Forcing | TaylorSeer | `interval` | 2.083 β **1.31x** | 2.583 β **1.59x** | 2.917 β **1.89x** | | |
| | Causal-Forcing | MotionCache | `thresh` | 0.7695 β **1.30x** | 1.2647 β **1.57x** | 2.8750 β **1.96x** | | |
| ### What a cached step actually costs | |
| Minima over ~1900 per-step samples per base, so these are solid: | |
| | method | cached step, as a fraction of a full step | ceiling at 14/28 forwards | | |
| |---|---|---| | |
| | TeaCache | 1.9% | 1.96x | | |
| | MotionCache | 2.0% | 1.96x | | |
| | FlowCache | 2.3% | 1.96x | | |
| | TaylorSeer | 5.6 - 5.8% | 1.89x | | |
| A full denoising forward costs 76-77 ms uncontended (1 chunk = 3 latent frames, | |
| 4680 tokens, Wan2.1-1.3B, H100). **2.0x is the arithmetic ceiling** of this | |
| schedule, and the achievable ceiling is ~1.96x for three of the methods. | |
| TaylorSeer pays more because its cached step is a Taylor *forecast*, not a copy: | |
| 30 layers x 3 module features have to be extrapolated and re-modulated. That is | |
| already with the `torch.compile` fusion upstream uses; without it the same step | |
| costs ~25% of a full one and the method tops out near 1.7x. | |
| ### Three findings | |
| 1. **The TeaCache-family indicator carries almost no signal on a 4-step distilled | |
| model.** Calibration over 252 (chunk, step) pairs: at a given denoising step the | |
| indicator's spread *across chunks* is a standard deviation of 0.0013 on a mean | |
| of 1.35 (0.1%), while the gaps *between* steps are 0.04-0.10. So the threshold | |
| is really choosing *which step index* to skip, not adapting per chunk. Worse, | |
| within a step the indicator is *negatively* correlated with the actual reuse | |
| error (r = -0.54 / -0.58 at steps 1 and 2); the +0.49 overall correlation is | |
| entirely a between-step effect, and a degree-4 fit only reaches R^2 = 0.26. | |
| 2. **That makes 1.6x hard to hit for the threshold methods and easy for the other | |
| two.** TeaCache's reachable points on Self-Forcing are essentially quantised to | |
| {1.0, 1.33, ~1.48, 2.0} β the requested 1.6x falls in a gap, and the closest | |
| honest point is 1.48x. FlowCache, deciding per frame group, does slightly better | |
| (1.51x on Self-Forcing, 1.61x on Causal-Forcing). TaylorSeer (fractional | |
| interval) and MotionCache (per token) land on 1.59x / 1.60x directly, because | |
| both have a genuinely continuous knob. | |
| 3. **Generalisation splits the methods the same way.** Re-running each setting on | |
| 64 prompts it was never tuned on, the compute fraction drifts by: | |
| | method | worst held-out drift | | |
| |---|---| | |
| | TaylorSeer | 0.000 | | |
| | MotionCache | 0.008 | | |
| | TeaCache | 0.079 | | |
| | FlowCache | 0.070 | | |
| TeaCache's and FlowCache's 1.3x and 1.6x settings sit inside a razor-thin band | |
| of the indicator distribution, so they do not transfer; their 2.0x settings | |
| (where everything skips regardless) transfer perfectly. TaylorSeer's schedule is | |
| content-independent by construction, and MotionCache's per-token thresholding | |
| averages over 4680 decisions per step instead of one. | |
| ### Artifacts for a quality comparison | |
| `videos/<base>_<method>_x<target>/NNNN.mp4` holds 7-10 clips per operating point, | |
| and `videos/<base>_baseline/NNNN.mp4` the matching all-full references (same | |
| prompts, same seed, indices align). The quality numbers are in `Experiment.md` | |
| (Extended-251 evaluation, next section). | |
| --- | |
| # Self-Forcing Extended-251 Full Evaluation | |
| VBench-8 + FFFF-relative pixel metrics + matched-prompt policy latency, following | |
| `self_forcing_extended_full_evaluation_protocol.md`, with the protocol's paths | |
| remapped to this machine. | |
| ## Path remapping | |
| | Protocol path | Here | | |
| |---|---| | |
| | `/data3/chenzhuo/anaconda3/envs/self_forcing` | `/mnt/local_nvme/zoubin/cz/envs/self_forcing` | | |
| | `$SELF_FORCING_REPO/.evaluation_env/vbench_site2` | absent -> `/mnt/local_nvme/zoubin/cz/projects/VBench` on `PYTHONPATH` | | |
| | `/data3/chenzhuo/.cache/vbench` | `/mnt/local_nvme/zoubin/cache/vbench` | | |
| | `prompts/vbench/all_dimension{,_extended}.txt` | `../Self-Forcing/prompts/vbench/` | | |
| | `assets/vbench8_extended_subset_mapping.json` | absent -> rebuilt by `eval/build_mapping.py` | | |
| ## Files | |
| ``` | |
| assets/vbench8_extended_subset_mapping.json the 251-row mapping (built + validated here) | |
| eval/build_mapping.py 946 -> 251 selection, with the protocol's validation gate | |
| eval/strategies.py FFFF + the 12 settled operating points, per base model | |
| eval/generate_eval.py video generation, latency split, pixel metrics | |
| eval/pixel_metrics.py PSNR / SSIM / LPIPS on decoded frames, pre-encode | |
| eval/run_vbench.py the eight dimensions, stock VBench metric code | |
| eval/aggregate.py normalize -> Quality/Semantic/Selected, latency, final tables | |
| eval/test_protocol.py the section 12 unit tests | |
| eval/run_generation.sh / run_vbench_all.sh / run_all_phases.sh 8-GPU drivers | |
| eval_out/ generated_videos/ per_prompt/ vbench/ summaries/ | |
| ``` | |
| ## Strategies | |
| 26 = 2 base models x (FFFF + 4 cache methods x 3 speed targets). FFFF is the | |
| all-full 4-step baseline and is each base model's own reference for both the | |
| pixel metrics and the latency speedup. Cache parameters are read back from | |
| `results/final_<base>*.json`, so the evaluation scores exactly the operating | |
| points the speed study settled on. | |
| ## Two environments, on purpose | |
| Generation runs in the **base `self_forcing` env** -- the same interpreter and | |
| package set the speed study used, so the evaluated videos come from exactly the | |
| benchmarked stack. The only addition to that env is `lpips`. | |
| VBench runs in the **`vbench_eval` venv**. It has to: the base env carries | |
| `torchao 0.17.0`, which is unimportable under `torch 2.5.1` (it references | |
| `torch.int1`), and `transformers.modeling_utils` pulls torchao in through its | |
| quantizer registry -- so `BertModel`, and therefore Tag2Text and the `scene` | |
| dimension, cannot load there. The venv pins `torchao==0.7.0`, | |
| `transformers==4.46.3` and adds `fairscale`. This is a deviation from protocol | |
| section 2.1 (one environment for everything); it does not touch generation. | |
| The venv itself needed repair first: a bad restore had flattened its `bin/python` | |
| symlinks into text files and stripped every executable bit. | |
| ## Protocol conformance | |
| Verified before the run: | |
| * mapping is 251 rows = 72 / 93 / 86, `global_index` unique, `suite_index` | |
| contiguous, all 86 scene rows keep `auxiliary_info`, and the 946 short prompts | |
| match `VBench_full_info.json` in order exactly (selection is by index -- two | |
| prompt texts are duplicated, so text matching is unsafe and is forbidden); | |
| * SHA256 recorded for the short prompts, extended prompts, `VBench_full_info.json` | |
| and the mapping; | |
| * videos are 81 frames, 480x832, 16 FPS, seed 0, one sample per prompt; | |
| * `prompt_en` handed to VBench is the **extended** prompt actually generated from; | |
| * each dimension scores only its own suite (72 / 93 / 86), driven by the mapping; | |
| * `policy_latency_ms` covers the denoise path only; the context/KV-cache DiT is | |
| timed separately into `excluded_context_kv_latency_ms` (~790 ms/video) and never | |
| added in; | |
| * pixel metrics run on the decoded RGB tensors before MP4 encoding, over all 81 | |
| frames, with FFFF-vs-itself pinned to 120 dB / 1.0 / 0.0; | |
| * `eval/test_protocol.py` covers the section 12 checklist and passes. | |