69.5 GB
49,395 files
Updated about 1 month ago
Name
Size
.claude
.git
.vscode
__pycache__
data_processing
evaluation
flexcombine-eval
results
tools
.gitignore70 Bytes
xet
README.md1.34 kB
xet
SETUP.md3.31 kB
xet
environment.yml486 Bytes
xet
requirements.txt465 Bytes
xet
test.mp4118 kB
xet
README.md

Benchmark construction

Dataset Source

https://huggingface.co/datasets/Fudan-FUXI/VIDGEN-1M

Dataset Scale Caption Length Key Features Best For
VidGen-1M 1M ~89 words Small but high-quality; coarse-to-fine filtering; strong temporal consistency and text-video alignment Clean T2V finetuning / low-resource baseline
Panda-70M 70M ~13 words Largest scale; diverse open-domain clips; captions selected from multiple teacher models Large-scale pretraining and diversity
Koala-36M 36M ~200 words Detailed structured captions; accurate temporal splitting; VTSS-based quality filtering Fine-grained prompt-video alignment

Types (single)

  • spatial

    • depth
    • pose
    • canny
  • subject

  • Style

  • inpainting / outpainting

Tools

Total size
69.5 GB
Files
49,395
Last updated
Aug 20
Pre-warmed CDN
US EU US EU

Contributors