LJHHHHHH/VideoX-Qwen-Datasets
Updated • 69 • 1
Model Weights Coming Soon
VideoX-Qwen is a unified model for general instruction-based video editing. It combines multimodal semantic conditioning with dense source-video latent guidance and supports object addition, removal, replacement, and attribute editing through natural-language instructions.
The model weights, inference code, configuration files, and usage instructions will be released soon.
VideoX-Qwen is trained with large-scale image and video editing supervision, including more than 1.2 million directional video-editing records constructed for this project.
The VideoX-Qwen dataset is available at:
The following resources will be provided upon release:
Please stay tuned for updates.