You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.
Log in or Sign Up to review the conditions and access this model content.
videoeditor-56b (35-expert scale-up of videoeditor-14b)
Same architecture as videoeditor-14b but every expert is replicated five times with small independent weight noise (7 sources x 5 copies = 35 experts per MoE layer, top-3), giving 55.93B parameters with the same ~8B active per token. Trained from the videoeditor-14b experts with expert-parallel/ZeRO-3 training on 24 H100s.
Training checkpoint: step 42,500.
Optimizer state: removed (weights only).
Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/videoeditor-56b/)
Package contents
The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).
| chunk | size | sha256 |
|---|---|---|
videoeditor-56b.tar.gz.part-00 |
21.47 GB | 4d8f89aada4d7856da34e153936c1567b8c32d5a5365083061f4bb0961578463 |
videoeditor-56b.tar.gz.part-01 |
21.47 GB | cd901823e901222946617d3a5406cbb2917c684eb63ec6ed1299da4e550a9296 |
videoeditor-56b.tar.gz.part-02 |
21.47 GB | 177a11f470a937f86b50aeae6c334ca61b8c14357ac8953bb53723a23cd4c1e4 |
videoeditor-56b.tar.gz.part-03 |
21.47 GB | acae09e788a927fa8db5cd6367d972c9b6b7307319b66e8de7cd5ba22012505d |
videoeditor-56b.tar.gz.part-04 |
3.79 GB | b098fb360347f3c8dcdc0d741a395d079f657af1af21db26521fc07ee6efd313 |
Total: 5 chunks, 89.69 GB compressed / 111.85 GB extracted.
Reassemble and verify:
huggingface-cli download Massyzs/videoeditor-56b --local-dir videoeditor-56b-package # after your access request is approved
cd videoeditor-56b-package
cat videoeditor-56b.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/ # creates <BASE_DIR>/ckpt/videoeditor-56b/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents) # optional check
MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).
Extracted files:
| file | size | sha256 |
|---|---|---|
videoeditor-56b/meta.json |
1059 B | 5d921b677850305e6edc5201972daebd59113dd2a8f11c10d9fa13bd6aa4592d |
videoeditor-56b/moe_v3_2_trainable.safetensors |
111.85 GB | 785987373973e5f84ce82d0bd0f1422911095e36fa56860204242157f81e4500 |
Base weights you also need
The package contains the full DiT (55.93B parameters, bf16, 111.9 GB); only the MLLM encoder (qwen/), the VAE and moe_expert_init.safetensors are needed in addition.
All frozen base components are published, already converted to the layout this code expects, in
Massyzs/kiwi-edit-5b-instruct-only-videoxfun:
huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun
See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.
Inference
CUDA_VISIBLE_DEVICES=0,1 python scripts/moe/videoeditor-56b/moe_v3_2_infer.py --gpus 2 --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/videoeditor-56b \
--src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"
The 56B DiT does not fit one 80 GB GPU: the released inference script splits the 30 blocks across two GPUs (blocks 0-14 on GPU 0, 15-29 + head on GPU 1) and needs about 120 GB of host RAM to build the model. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, segmented text CFG (scale 3 on the first 60% of the steps). Output: <OUT_DIR>/<name>.mp4 plus lossless PNG frames and expert-routing maps.
Architecture
As videoeditor-14b, with 35 experts per MoE layer (expert j is a noised copy of source j//5: main, style, add, remove, replace, background, direct_finetune). Full state dict: 4,663 tensors, 111.85 GB bf16.
Training data
OpenVE + ReCo single-instruction edits mixed with multi-instruction data (goku composite + CoinVE, x10 repeat); 42,500 steps, effective batch 120 (5 x 24 GPUs).
License
The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.
- Downloads last month
- -
Model tree for Massyzs/videoeditor-56b
Base model
linyq/kiwi-edit-5b-instruct-only-diffusers