Video-to-Video
PyTorch
FlashRender / README.md
byeongjun-park's picture
Add stage 1/2/3 checkpoints, config.json, and model card
a9a7638 verified
|
Raw History Blame Contribute Delete
1.53 kB
---
license: apache-2.0
datasets:
- KlingTeam/MultiCamVideo-Dataset
- KlingTeam/SynCamVideo-Dataset
base_model:
- alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera
base_model_relation: finetune
pipeline_tag: video-to-video
tags:
- pytorch
---
This model was presented in [FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow](https://huggingface.co/papers/2609.03563).
Please refer to the [Github](https://github.com/byeongjun-park/FlashRender) README for usage.
## Checkpoints
| Stage | File | Sampling |
|:---:|---|---|
| 1 | `epoch_20.safetensors` | 50 steps, `inference_mode=mul` |
| 2 | `meanflow_epoch_20.safetensors` | 4 steps, `inference_mode=any` |
| 3 | `onpolicy_epoch_5.safetensors` | 4 steps, `inference_mode=any` |
**`onpolicy_epoch_5.safetensors` is the final model** and the default in `configs/base.yaml`.
Each file stores only the **trainable parameter subset**, not a stand-alone diffusers/transformers
checkpoint. Weights are `bfloat16`, except the relative-pose tokens and the per-block RoPE phase MLPs
(`rel_pose_{src,tgt}_token`, `self_attn.rope_phase_{qk,vo}`), which are stored in `float32`.
They are loaded with `strict=False` onto the base Wan DiT after it is patched by
`flashrender_utils.model_utils.adjust_to_FlashRender`, so use the code repository:
```python
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="byeongjun-park/FlashRender",
local_dir="models/checkpoints",
allow_patterns=["*.safetensors", "config.json"],
)
```