Video-to-Video
PyTorch
File size: 1,526 Bytes
8bc1a5b
 
a9a7638
 
 
 
 
 
 
 
 
8bc1a5b
a9a7638
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
---
license: apache-2.0
datasets:
- KlingTeam/MultiCamVideo-Dataset
- KlingTeam/SynCamVideo-Dataset
base_model:
- alibaba-pai/Wan2.1-Fun-V1.1-1.3B-Control-Camera
base_model_relation: finetune
pipeline_tag: video-to-video
tags:
- pytorch
---

This model was presented in [FlashRender: Few-Step Generative Rendering via Camera-Controlled Video MeanFlow](https://huggingface.co/papers/2609.03563).

Please refer to the [Github](https://github.com/byeongjun-park/FlashRender) README for usage.

## Checkpoints

| Stage | File | Sampling |
|:---:|---|---|
| 1 | `epoch_20.safetensors` | 50 steps, `inference_mode=mul` |
| 2 | `meanflow_epoch_20.safetensors` | 4 steps, `inference_mode=any` |
| 3 | `onpolicy_epoch_5.safetensors` | 4 steps, `inference_mode=any` |

**`onpolicy_epoch_5.safetensors` is the final model** and the default in `configs/base.yaml`.

Each file stores only the **trainable parameter subset**, not a stand-alone diffusers/transformers
checkpoint. Weights are `bfloat16`, except the relative-pose tokens and the per-block RoPE phase MLPs
(`rel_pose_{src,tgt}_token`, `self_attn.rope_phase_{qk,vo}`), which are stored in `float32`.
They are loaded with `strict=False` onto the base Wan DiT after it is patched by
`flashrender_utils.model_utils.adjust_to_FlashRender`, so use the code repository:

```python
from huggingface_hub import snapshot_download

snapshot_download(
    repo_id="byeongjun-park/FlashRender",
    local_dir="models/checkpoints",
    allow_patterns=["*.safetensors", "config.json"],
)
```