# Cache-method comparison on 4-step block-causal video DiTs Four training-free caching methods (TeaCache, TaylorSeer, FlowCache, MotionCache) implemented on two 4-step autoregressive video base models (Self-Forcing and Causal-Forcing), swept to matched denoise-DiT speedups. ## Layout ``` repos/ the six upstream clones (reference implementations + base models) Self-Forcing/ base model 1 (guandeh17/Self-Forcing) Causal-Forcing/ base model 2 (thu-ml/Causal-Forcing) TeaCache/ ali-vilab/TeaCache TaylorSeer/ Shenyi-Z/TaylorSeer FlowCache/ mikeallen39/FlowCache (ICLR 2026) MotionCache/ MAC-AutoML/MotionCache (ICML 2026) cachelib/ the ports -- one implementation, both base models methods.py the four cache methods + baseline + calibration probe selective.py token-subset forward for the KV-cached causal DiT patch.py routes CausalWanModel._forward_inference through a method runner.py chunked denoising loop with cache hooks and denoise-only timing harness.py model loading, prompt loading, paired A/B measurement calibrate.py fits the TeaCache rescale polynomial for a base model sweep.py finds the parameter that hits a target speedup finalize.py re-targets on the timed prompts + held-out check (authoritative) overhead_bench.py per-step cost of a cached vs a full step (contention-robust) retime.py plain stopwatch re-measurement of the settled points verify.py re-measures operating points on held-out prompts run_compare.py single-configuration CLI summarize.py collects everything into one table results/ JSON output, plus SUMMARY.md / SUMMARY.csv videos/ generated mp4s per operating point, plus *_baseline references ``` Pipeline order: `calibrate.py` -> `sweep.py` -> `finalize.py` -> `overhead_bench.py` -> `retime.py` -> `summarize.py`. Weights are symlinked from the pre-existing checkouts rather than re-downloaded: | base | checkpoint | source | |---|---|---| | self_forcing | `checkpoints/self_forcing_dmd.pt` (`generator_ema`) | `../Self-Forcing` | | causal_forcing | `checkpoints/chunkwise/causal_forcing.pt` (`generator`) | `../Causal-Forcing` | Both also symlink `wan_models/` (Wan2.1-T2V-1.3B backbone, VAE, UMT5 text encoder). ## The setting Both base models are the same architecture: a block-causal Wan2.1-1.3B DiT that generates a video chunk at a time. With `num_frame_per_block=3` and 21 latent frames, a video is **7 chunks x 4 denoising steps = 28 DiT forwards**, plus one untimed KV-cache refresh pass per chunk (`context_noise`) that rewrites the chunk's KV entries from the clean latent. The experiment fixes the schedule to **`F ? ? F`**: step 0 and step 3 always run the full DiT, steps 1 and 2 are the ones a cache method may skip. This is equivalent to the upstream `ret_steps=1` / `cutoff_steps=num_steps-1` guards. **Speedup is measured on denoise-DiT time only** — the sum of the 28 denoising forwards, timed with CUDA events. Text encoding, VAE decode and the per-chunk KV-refresh pass are excluded. That makes 2.0x the arithmetic ceiling (14 of 28 forwards), which is why the 2.0x target is exactly "skip both middle steps everywhere". ## How each method was ported All four hook the same place: the 30-block loop inside `CausalWanModel._forward_inference`. The DiT preamble and the head/unpatchify tail always run. | method | decision granularity | what a skipped step reuses | knob | |---|---|---|---| | TeaCache | whole forward | the 30-block residual `x_out - x_in` from the last computed step | `thresh` | | FlowCache | per frame group inside the chunk | the same residual, banked per group | `thresh`, `group_size` | | TaylorSeer | whole forward | a **Taylor forecast** of each block's self-attn / cross-attn / FFN output | `interval`, `max_order` | | MotionCache | per token, motion-weighted | cached rows for un-selected tokens; selected tokens are recomputed | `thresh`, `weight_norm` | Skipping the block stack leaves the chunk's KV-cache entries holding the previous computed step's keys/values. That is safe here: every later step of the chunk overwrites the same slots, and the per-chunk KV-refresh pass rewrites them from the clean latent before the next chunk reads them. ### Selective (partial-token) forward FlowCache and MotionCache recompute only part of the chunk. On a KV-cached causal model that means, per self-attention layer: compute q/k/v for the selected rows only, scatter their fresh k/v into the slots the chunk owns (leaving unselected slots holding the previous step's k/v), and attend the selected queries against the full cache. This mirrors MotionCache's `forward_selective`, moved from SkyReels' persistent buffers onto Self-Forcing's rolling `kv_cache`. Validated: with *all* tokens selected the selective path is **bit-exact** against the stock full forward (max abs latent difference 0.0). The same check passes for TaylorSeer's recording path at `interval=1`. ### Deviations from upstream, and why 1. **Indicator.** `TeaCache4Wan2.1` derives its indicator from `e0` (the timestep modulation) alone. That is fine for bidirectional Wan, where one forward covers the whole video at a single timestep, but it *degenerates* here: every chunk runs the same four timesteps, so an `e0` indicator is bit-identical across chunks and the threshold cannot adapt to content at all. The default is therefore TeaCache's original definition — the relative L1 distance of the **timestep-modulated noisy input**, the tensor block 0 feeds to its self-attention, which is what TeaCache uses on HunyuanVideo and FLUX. `--indicator e0` restores the literal Wan port. 2. **Rescale polynomial.** TeaCache's shipped 4th-degree coefficients were fitted for a different model, schedule and indicator, so `calibrate.py` reruns TeaCache's own fitting procedure here. The fit is **degree 1, not 4**: on this model the input distance only spans ~[1.19, 1.36], where a degree-4 fit is ill-conditioned (coefficients ~10^3 with alternating signs) and MotionCache, which evaluates the polynomial *per token*, would extrapolate it far outside the fitted domain. Degree 1 is monotone and safe. All degrees are recorded in the coefficient JSON. 3. **FlowCache granularity.** FlowCache's contribution is that the units covered by one forward denoise at different rates and need independent policies. In SkyReels-V2 those units are the chunks of the diffusion-forcing window; here a forward covers exactly one chunk, so the corresponding units are the frame groups inside it (`group_size=1`, i.e. 3 groups). Its KV-cache-compression component is not ported: Self-Forcing already bounds its KV cache by local attention, so there is no growing cache to compress. 4. **TaylorSeer fractional interval.** TaylorSeer's schedule is a uniform integer refresh interval, which on an `F ? ? F` schedule can only produce 1.0x, 1.33x and 2.0x. `interval` is accepted as a float and realised as a deterministic per-chunk alternation between the two bracketing integer intervals (Bresenham-style), giving the requested *average* interval. This keeps TaylorSeer's content-independent schedule rather than bolting a threshold onto it. 5. **Fused forecast.** TaylorSeer's cached step is ~90 bandwidth-bound elementwise ops per forward. Upstream puts `@torch.compile` on the equivalent `wan_attention_cache_forward`; the port does the same. Without it the forecast ate most of the saving (1.72x measured at a 2.0x FLOPs setting); with it, it does not. Set `CACHELIB_NO_COMPILE=1` to disable. ## Scope change (Sep 4) FlowCache is frozen: its finished operating points stay in the tables, nothing new is generated for it. The continuing study is **three methods** (TeaCache, TaylorSeer, MotionCache) x **three base models** (Self-Forcing, Causal-Forcing, HY-WorldPlay) x **three points**: `FxxF` at ~1.3x, `FxxF` at 1.7-1.8x, and `Fxxx` at 2.7-3.0x -- all targeted on **measured** wall-clock speedup, not on the compute fraction. The reason is MotionCache: matching compute at 2.0x under `FxxF` forces its per-token selection to pick *zero* tokens, which makes it bit-identical to TeaCache and reproduces nothing. Operating points are now chosen as the fastest setting that still keeps the token-level decision active (>=25% of cacheable steps partial, >=5% of tokens selected on those steps); the activity statistics are recorded per video in ``cache_diagnostics``. ## Schedules beyond `FxxF` Every entry point takes `--schedule` (`sweep.py`, `finalize.py`, `run_compare.py`; the evaluation reads it from the `final_*.json` rows). `F` marks a step that always runs the full DiT, `x` one the method may serve from cache: `FxxF` is the study above and the default; `FFxx` gives TaylorSeer two computed points before its first forecast so the first-order term is actually exercised; `Fxxx` lifts the cool-down guard and raises the arithmetic ceiling to 4x. Operating points for `FFxx` (TaylorSeer, 1.3/1.6/2.0x) and `Fxxx` (the other three, ~3x) were searched on 3 prompts (`results/sweep_*_{FFxx,Fxxx}.json`, collected in `results/final__newsched.json`) and evaluated on Extended-251. **All results, speed and quality, are collected in `Experiment.md`** (regenerate with `python experiment_md.py`). The headline of the second round: caching the *last* denoising step -- which every `Fxxx` point and every `FFxx` point at >= 1.6x does -- costs 15-25 VBench points on both base models, while `FxxF` at the same compute costs 1-2. On a 4-step distilled model the final step is not optional. ## Measurement protocol This is a **shared node** and it was busy throughout. Other tenants' jobs move absolute timings by 30%+ and drift them *within* a single sweep, so a baseline measured once at the start silently inflates every later configuration. Four things follow, and they are the difference between numbers worth reading and numbers that are noise: 1. **Search on compute, not on time.** The parameter search runs on the compute fraction (how many of the 28 DiT forwards were actually computed), which is exactly deterministic given the prompt set and immune to contention. Targets are aimed at the compute budget directly, which also puts all four methods on identical compute — the right basis for a later quality comparison. Correcting the target by a measured overhead ratio was tried and abandoned: that ratio is itself a timing measurement, one probe returned 1.110 (physically impossible), and it dragged a whole method's operating points below target. 2. **Paired A/B.** Baseline and method run back to back on the same prompt, so drift between them is seconds rather than minutes. 3. **Minimum, not median, as the headline.** Contention is one-sided — a neighbour can only make a run slower — so the fastest observed run is the best estimate of the uncontended time, and the headline speedup is `min(baseline) / min(method)` over the pairs. The paired median is reported alongside but is biased *upward*, because the baseline run is longer and so more exposed to being hit by a spike; that is why it sometimes exceeded a method's arithmetic ceiling. 4. **Compute fraction from the timed runs.** A threshold tuned on one prompt subset and timed on another silently disagrees whenever it sits near a decision boundary. `finalize.py` reports the compute fraction of the very runs it timed, and re-tunes on those prompts if it misses. `finalize.py` also re-checks every point on a disjoint prompt slice, so a threshold that only hit its target by sitting inside a narrow band of the indicator distribution shows up as held-out drift. `verify.py` does the same check standalone against a sweep file. ## Reproducing ```bash PY=/local/zoubin/cz/envs/self_forcing/bin/python # 1. fit the rescale polynomial (once per base model) CUDA_VISIBLE_DEVICES=0 $PY calibrate.py --base self_forcing \ --num-prompts 12 --out results/coeff_self_forcing.json # 2. sweep one method to the three target speedups CUDA_VISIBLE_DEVICES=0 $PY sweep.py --base self_forcing --method teacache \ --coefficients results/coeff_self_forcing.json \ --targets 1.3,1.6,2.0 --save-video-dir videos \ --out results/sweep_self_forcing_teacache.json # 3. check the settings transfer to unseen prompts CUDA_VISIBLE_DEVICES=0 $PY verify.py \ --sweep results/sweep_self_forcing_teacache.json --prompt-offset 64 \ --out results/verify_self_forcing_teacache.json # 4. collect everything $PY summarize.py ``` `run_compare.py` runs a single configuration if you just want one number: ```bash CUDA_VISIBLE_DEVICES=0 $PY run_compare.py --base causal_forcing \ --method taylorseer --interval 2.5 --num-prompts 8 ``` ## Environment note **Sep 3:** the machine was restored again -- `/mnt/local_nvme` became `/local`, both environments lost their executable bits a second time (1632 files restored), the `vbench_eval` venv's `pyvenv.cfg`/interpreter links pointed at the old path (now symlinks into `/local/.../self_forcing`), and the caches moved to `/local/zoubin/cz/.cache/`. Paths in `eval/*.sh`, `eval/aggregate.py` and the per-prompt records were updated accordingly. The node is shared: other tenants' launches have repeatedly killed GPU processes of ours, so long runs go through `eval/retime_until_clean.sh`-style loops that only use quiet GPUs and resume. The conda environment at `/mnt/local_nvme/zoubin/cz/envs/self_forcing` had lost every executable bit (a bad restore left all files `rw-rw-r--`), so nothing in it could run — including Triton's bundled `ptxas`, which broke `torch.compile`. The exec bits were restored on files that are ELF binaries or start with `#!` (1772 files). No package contents were changed. ## Results Full tables: `results/SUMMARY.md` (regenerate with `python summarize.py`), machine readable in `results/SUMMARY.csv` and the per-stage JSONs. **Headline speedup** is the denoise-DiT speedup implied by the measured compute fraction and the measured per-step cost model (see Measurement protocol). The `stopwatch` column is the plain paired whole-video ratio; it agrees within its (large) spread but is not the number to quote from this node. ### Operating points found | base | method | knob | 1.3x | 1.6x | 2.0x | |---|---|---|---|---|---| | Self-Forcing | TeaCache | `thresh` | 0.6348 → **1.30x** | 1.2507 → **1.48x** | 1.2917 → **1.96x** | | Self-Forcing | FlowCache | `thresh` | 0.6360 → **1.30x** | 1.2500 → **1.51x** | 1.2917 → **1.96x** | | Self-Forcing | TaylorSeer | `interval` | 2.083 → **1.31x** | 2.583 → **1.59x** | 2.917 → **1.89x** | | Self-Forcing | MotionCache | `thresh` | 0.7560 → **1.29x** | 1.2461 → **1.60x** | 2.5000 → **1.96x** | | Causal-Forcing | TeaCache | `thresh` | 0.6491 → **1.30x** | 1.2760 → **1.51x** | 1.2917 → **1.96x** | | Causal-Forcing | FlowCache | `thresh` | 0.6488 → **1.26x** | 1.2760 → **1.61x** | 1.2917 → **1.95x** | | Causal-Forcing | TaylorSeer | `interval` | 2.083 → **1.31x** | 2.583 → **1.59x** | 2.917 → **1.89x** | | Causal-Forcing | MotionCache | `thresh` | 0.7695 → **1.30x** | 1.2647 → **1.57x** | 2.8750 → **1.96x** | ### What a cached step actually costs Minima over ~1900 per-step samples per base, so these are solid: | method | cached step, as a fraction of a full step | ceiling at 14/28 forwards | |---|---|---| | TeaCache | 1.9% | 1.96x | | MotionCache | 2.0% | 1.96x | | FlowCache | 2.3% | 1.96x | | TaylorSeer | 5.6 - 5.8% | 1.89x | A full denoising forward costs 76-77 ms uncontended (1 chunk = 3 latent frames, 4680 tokens, Wan2.1-1.3B, H100). **2.0x is the arithmetic ceiling** of this schedule, and the achievable ceiling is ~1.96x for three of the methods. TaylorSeer pays more because its cached step is a Taylor *forecast*, not a copy: 30 layers x 3 module features have to be extrapolated and re-modulated. That is already with the `torch.compile` fusion upstream uses; without it the same step costs ~25% of a full one and the method tops out near 1.7x. ### Three findings 1. **The TeaCache-family indicator carries almost no signal on a 4-step distilled model.** Calibration over 252 (chunk, step) pairs: at a given denoising step the indicator's spread *across chunks* is a standard deviation of 0.0013 on a mean of 1.35 (0.1%), while the gaps *between* steps are 0.04-0.10. So the threshold is really choosing *which step index* to skip, not adapting per chunk. Worse, within a step the indicator is *negatively* correlated with the actual reuse error (r = -0.54 / -0.58 at steps 1 and 2); the +0.49 overall correlation is entirely a between-step effect, and a degree-4 fit only reaches R^2 = 0.26. 2. **That makes 1.6x hard to hit for the threshold methods and easy for the other two.** TeaCache's reachable points on Self-Forcing are essentially quantised to {1.0, 1.33, ~1.48, 2.0} — the requested 1.6x falls in a gap, and the closest honest point is 1.48x. FlowCache, deciding per frame group, does slightly better (1.51x on Self-Forcing, 1.61x on Causal-Forcing). TaylorSeer (fractional interval) and MotionCache (per token) land on 1.59x / 1.60x directly, because both have a genuinely continuous knob. 3. **Generalisation splits the methods the same way.** Re-running each setting on 64 prompts it was never tuned on, the compute fraction drifts by: | method | worst held-out drift | |---|---| | TaylorSeer | 0.000 | | MotionCache | 0.008 | | TeaCache | 0.079 | | FlowCache | 0.070 | TeaCache's and FlowCache's 1.3x and 1.6x settings sit inside a razor-thin band of the indicator distribution, so they do not transfer; their 2.0x settings (where everything skips regardless) transfer perfectly. TaylorSeer's schedule is content-independent by construction, and MotionCache's per-token thresholding averages over 4680 decisions per step instead of one. ### Artifacts for a quality comparison `videos/__x/NNNN.mp4` holds 7-10 clips per operating point, and `videos/_baseline/NNNN.mp4` the matching all-full references (same prompts, same seed, indices align). The quality numbers are in `Experiment.md` (Extended-251 evaluation, next section). --- # Self-Forcing Extended-251 Full Evaluation VBench-8 + FFFF-relative pixel metrics + matched-prompt policy latency, following `self_forcing_extended_full_evaluation_protocol.md`, with the protocol's paths remapped to this machine. ## Path remapping | Protocol path | Here | |---|---| | `/data3/chenzhuo/anaconda3/envs/self_forcing` | `/mnt/local_nvme/zoubin/cz/envs/self_forcing` | | `$SELF_FORCING_REPO/.evaluation_env/vbench_site2` | absent -> `/mnt/local_nvme/zoubin/cz/projects/VBench` on `PYTHONPATH` | | `/data3/chenzhuo/.cache/vbench` | `/mnt/local_nvme/zoubin/cache/vbench` | | `prompts/vbench/all_dimension{,_extended}.txt` | `../Self-Forcing/prompts/vbench/` | | `assets/vbench8_extended_subset_mapping.json` | absent -> rebuilt by `eval/build_mapping.py` | ## Files ``` assets/vbench8_extended_subset_mapping.json the 251-row mapping (built + validated here) eval/build_mapping.py 946 -> 251 selection, with the protocol's validation gate eval/strategies.py FFFF + the 12 settled operating points, per base model eval/generate_eval.py video generation, latency split, pixel metrics eval/pixel_metrics.py PSNR / SSIM / LPIPS on decoded frames, pre-encode eval/run_vbench.py the eight dimensions, stock VBench metric code eval/aggregate.py normalize -> Quality/Semantic/Selected, latency, final tables eval/test_protocol.py the section 12 unit tests eval/run_generation.sh / run_vbench_all.sh / run_all_phases.sh 8-GPU drivers eval_out/ generated_videos/ per_prompt/ vbench/ summaries/ ``` ## Strategies 26 = 2 base models x (FFFF + 4 cache methods x 3 speed targets). FFFF is the all-full 4-step baseline and is each base model's own reference for both the pixel metrics and the latency speedup. Cache parameters are read back from `results/final_*.json`, so the evaluation scores exactly the operating points the speed study settled on. ## Two environments, on purpose Generation runs in the **base `self_forcing` env** -- the same interpreter and package set the speed study used, so the evaluated videos come from exactly the benchmarked stack. The only addition to that env is `lpips`. VBench runs in the **`vbench_eval` venv**. It has to: the base env carries `torchao 0.17.0`, which is unimportable under `torch 2.5.1` (it references `torch.int1`), and `transformers.modeling_utils` pulls torchao in through its quantizer registry -- so `BertModel`, and therefore Tag2Text and the `scene` dimension, cannot load there. The venv pins `torchao==0.7.0`, `transformers==4.46.3` and adds `fairscale`. This is a deviation from protocol section 2.1 (one environment for everything); it does not touch generation. The venv itself needed repair first: a bad restore had flattened its `bin/python` symlinks into text files and stripped every executable bit. ## Protocol conformance Verified before the run: * mapping is 251 rows = 72 / 93 / 86, `global_index` unique, `suite_index` contiguous, all 86 scene rows keep `auxiliary_info`, and the 946 short prompts match `VBench_full_info.json` in order exactly (selection is by index -- two prompt texts are duplicated, so text matching is unsafe and is forbidden); * SHA256 recorded for the short prompts, extended prompts, `VBench_full_info.json` and the mapping; * videos are 81 frames, 480x832, 16 FPS, seed 0, one sample per prompt; * `prompt_en` handed to VBench is the **extended** prompt actually generated from; * each dimension scores only its own suite (72 / 93 / 86), driven by the mapping; * `policy_latency_ms` covers the denoise path only; the context/KV-cache DiT is timed separately into `excluded_context_kv_latency_ms` (~790 ms/video) and never added in; * pixel metrics run on the decoded RGB tensors before MP4 encoding, over all 81 frames, with FFFF-vs-itself pinned to 120 dB / 1.0 / 0.0; * `eval/test_protocol.py` covers the section 12 checklist and passes.