VideoX-Qwen

Model Weights Coming Soon

VideoX-Qwen is a unified model for general instruction-based video editing. It combines multimodal semantic conditioning with dense source-video latent guidance and supports object addition, removal, replacement, and attribute editing through natural-language instructions.

The model weights, inference code, configuration files, and usage instructions will be released soon.

Dataset

VideoX-Qwen is trained with large-scale image and video editing supervision, including more than 1.2 million directional video-editing records constructed for this project.

The VideoX-Qwen dataset is available at:

VideoX-Qwen-Datasets

Release Plan

The following resources will be provided upon release:

  • Model weights
  • Inference code
  • Configuration files
  • Installation and usage instructions
  • Example editing results

Please stay tuned for updates.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train LJHHHHHH/VideoX-Qwen