Buckets:
69.5 GB
49,395 files
Updated about 1 month ago
Ctrl+K
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| .claude | 1 items | ||
| .git | 279 items | ||
| .vscode | 1 items | ||
| __pycache__ | 2 items | ||
| data_processing | 9 items | ||
| evaluation | 4,334 items | ||
| flexcombine-eval | 1 items | ||
| results | 35,507 items | ||
| tools | 9,255 items | ||
| .gitignore | 70 Bytes xet | dec5777f | |
| README.md | 1.34 kB xet | 776ba312 | |
| SETUP.md | 3.31 kB xet | 2497ca07 | |
| environment.yml | 486 Bytes xet | 5e4f51fe | |
| requirements.txt | 465 Bytes xet | fba044aa | |
| test.mp4 | 118 kB xet | ca5524e2 |
Benchmark construction
Dataset Source
https://huggingface.co/datasets/Fudan-FUXI/VIDGEN-1M
| Dataset | Scale | Caption Length | Key Features | Best For |
|---|---|---|---|---|
| VidGen-1M | 1M | ~89 words | Small but high-quality; coarse-to-fine filtering; strong temporal consistency and text-video alignment | Clean T2V finetuning / low-resource baseline |
| Panda-70M | 70M | ~13 words | Largest scale; diverse open-domain clips; captions selected from multiple teacher models | Large-scale pretraining and diversity |
| Koala-36M | 36M | ~200 words | Detailed structured captions; accurate temporal splitting; VTSS-based quality filtering | Fine-grained prompt-video alignment |
Types (single)
spatial
- depth
- pose
- canny
subject
Style
inpainting / outpainting
Tools
- Total size
- 69.5 GB
- Files
- 49,395
- Last updated
- Aug 20
- Pre-warmed CDN
- US EU US EU