You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.

Log in or Sign Up to review the conditions and access this model content.

videoeditor-56b (35-expert scale-up of videoeditor-14b)

Same architecture as videoeditor-14b but every expert is replicated five times with small independent weight noise (7 sources x 5 copies = 35 experts per MoE layer, top-3), giving 55.93B parameters with the same ~8B active per token. Trained from the videoeditor-14b experts with expert-parallel/ZeRO-3 training on 24 H100s.

Training checkpoint: step 42,500. Optimizer state: removed (weights only). Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/videoeditor-56b/)

Package contents

The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).

chunk size sha256
videoeditor-56b.tar.gz.part-00 21.47 GB 4d8f89aada4d7856da34e153936c1567b8c32d5a5365083061f4bb0961578463
videoeditor-56b.tar.gz.part-01 21.47 GB cd901823e901222946617d3a5406cbb2917c684eb63ec6ed1299da4e550a9296
videoeditor-56b.tar.gz.part-02 21.47 GB 177a11f470a937f86b50aeae6c334ca61b8c14357ac8953bb53723a23cd4c1e4
videoeditor-56b.tar.gz.part-03 21.47 GB acae09e788a927fa8db5cd6367d972c9b6b7307319b66e8de7cd5ba22012505d
videoeditor-56b.tar.gz.part-04 3.79 GB b098fb360347f3c8dcdc0d741a395d079f657af1af21db26521fc07ee6efd313

Total: 5 chunks, 89.69 GB compressed / 111.85 GB extracted.

Reassemble and verify:

huggingface-cli download Massyzs/videoeditor-56b --local-dir videoeditor-56b-package      # after your access request is approved
cd videoeditor-56b-package
cat videoeditor-56b.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/                # creates <BASE_DIR>/ckpt/videoeditor-56b/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents)   # optional check

MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).

Extracted files:

file size sha256
videoeditor-56b/meta.json 1059 B 5d921b677850305e6edc5201972daebd59113dd2a8f11c10d9fa13bd6aa4592d
videoeditor-56b/moe_v3_2_trainable.safetensors 111.85 GB 785987373973e5f84ce82d0bd0f1422911095e36fa56860204242157f81e4500

Base weights you also need

The package contains the full DiT (55.93B parameters, bf16, 111.9 GB); only the MLLM encoder (qwen/), the VAE and moe_expert_init.safetensors are needed in addition.

All frozen base components are published, already converted to the layout this code expects, in Massyzs/kiwi-edit-5b-instruct-only-videoxfun:

huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun

See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.

Inference

CUDA_VISIBLE_DEVICES=0,1 python scripts/moe/videoeditor-56b/moe_v3_2_infer.py --gpus 2 --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/videoeditor-56b \
    --src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"

The 56B DiT does not fit one 80 GB GPU: the released inference script splits the 30 blocks across two GPUs (blocks 0-14 on GPU 0, 15-29 + head on GPU 1) and needs about 120 GB of host RAM to build the model. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, segmented text CFG (scale 3 on the first 60% of the steps). Output: <OUT_DIR>/<name>.mp4 plus lossless PNG frames and expert-routing maps.

Architecture

As videoeditor-14b, with 35 experts per MoE layer (expert j is a noised copy of source j//5: main, style, add, remove, replace, background, direct_finetune). Full state dict: 4,663 tensors, 111.85 GB bf16.

Training data

OpenVE + ReCo single-instruction edits mixed with multi-instruction data (goku composite + CoinVE, x10 repeat); 42,500 steps, effective batch 120 (5 x 24 GPUs).

License

The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Massyzs/videoeditor-56b

Finetuned
(5)
this model