You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.

Log in or Sign Up to review the conditions and access this model content.

explicit-moe (Kiwi-Edit 5B + explicit task experts)

An instruction-based video-editing model that keeps the Kiwi-Edit 5B DiT frozen and adds five explicit task experts (style, add, remove, replace, background), each a LoRA (rank 128 on q/k/v/o and both FFN projections) fine-tuned on one edit category. A router head assigns every sub-instruction to one expert (argmax over the five categories); each expert runs as a separate DiT stream, and 14 trainable fusion attention layers after DiT blocks 15-28 let the base stream attend to the expert streams.

Training checkpoint: step 28,200. Optimizer state: removed (weights only). Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/explicit-moe/)

Package contents

The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).

chunk size sha256
explicit-moe.tar.gz.part-00 8.81 GB a3732ceeb496e294b7ead2625ae5fa6eb12a0be24a5356d3a0988e392e8c152d

Total: 1 chunks, 8.81 GB compressed / 9.67 GB extracted.

Reassemble and verify:

huggingface-cli download Massyzs/explicit-moe --local-dir explicit-moe-package      # after your access request is approved
cd explicit-moe-package
cat explicit-moe.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/                # creates <BASE_DIR>/ckpt/explicit-moe/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents)   # optional check

MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).

Extracted files:

file size sha256
explicit-moe/fusion_moe_v2_absorb.safetensors 1.06 GB ec6ff554367b337222e7e1f9ce9ca8bf094d584958a71fdf67fd6b7b757b6394
explicit-moe/lora_experts/add/dit_query_router.safetensors 1.29 GB e89781f56ce9ae1b6378cc1a00b8ae91cf0d89c17e39e1e3fdf06c66b28fb74b
explicit-moe/lora_experts/background/dit_query_router.safetensors 1.29 GB c7ed8394bd3493168130186c21f189803131d188e968bf745ef160f3b4753fac
explicit-moe/lora_experts/remove/dit_query_router.safetensors 1.29 GB 58a8936c1f208e55266f624d97306008698eaf57de3e2ed4a73796190364a011
explicit-moe/lora_experts/replace/dit_query_router.safetensors 1.29 GB 88406dc90ce0a79b591cbcb2ce0f9762f1e97e4c5a421c1d21133bd9d88db0c1
explicit-moe/lora_experts/style/dit_query_router.safetensors 1.29 GB d5f03ea7aea848e5d108b5dc596d135400913670a0ebb4a16bba109e6c4cfde8
explicit-moe/meta.json 1038 B d407255d560faafd42249393a4f3f6dbcbd3093f0a245d6b6da978349950645f
explicit-moe/moe_v2_state.pt 2.13 GB 07929c4ed8ad28dc1560a5508ed7e6ca20d79f12e1664d860ab4cccf13fed0b0
explicit-moe/router_head_finetuned.pt 16.8 MB fc99904f4870e610c1b758a1638a0b735e095f824d9a9cc76b928bc504a3a902
explicit-moe/router_head_online.pt 16.8 MB 1695f5e0defe34a95ec9dc4bfa0bd9e41b08d602aed8bc908f66bce42f41471b

Base weights you also need

The DiT base weights are required: the LoRA experts and fusion layers are applied on top of the frozen Kiwi-Edit 5B DiT.

All frozen base components are published, already converted to the layout this code expects, in Massyzs/kiwi-edit-5b-instruct-only-videoxfun:

huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun

See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.

Inference

python scripts/moe/explicit-moe/moe_v2_infer.py --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/explicit-moe \
    --src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"

Single 80 GB GPU. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, text CFG 3 on all steps. Output: <OUT_DIR>/<name>.mp4 plus lossless PNG frames in <OUT_DIR>/frames/<name>/. moe_v2_state.pt (fp32 trainable tensors, optimizer removed) is what the loader reads; fusion_moe_v2_absorb.safetensors and router_head_finetuned.pt contain the same tensors in bf16 / as a state_dict for convenience. router_head_online.pt supplies only the router-head architecture fields.

Architecture

Frozen Kiwi-Edit DiT (30 blocks, dim 3072) + 5 LoRA experts (0.32B parameters each) + router head (self-attention over the MLLM instruction-token features, 5-way classifier; argmax selects the expert of each sub-instruction) + 14 fusion attention layers (0.53B parameters) after blocks 15-28 in which the base stream attends to itself and to the expert streams. Trainable parameters: 188 tensors, 2.13 GB fp32 (fusion layers + router head).

Training data

OpenVE single-instruction edits and multi-instruction data (goku composite + CoinVE), ~2 epochs, effective batch 32 (2 x 16 GPUs).

License

The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Massyzs/explicit-moe

Finetuned
(5)
this model