You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.
Log in or Sign Up to review the conditions and access this model content.
mlp-moe (Kiwi-Edit 5B with a 12-expert FFN mixture-of-experts)
The Kiwi-Edit 5B DiT with the FFN of 14 of its 30 blocks (blocks 1,3,...,27) replaced by a DeepSeek-style mixture of 12 full-weight experts (top-3 routing, no shared expert, expert-level balance loss). Experts are initialised from the five category LoRA experts and a direct full-parameter fine-tune (two noised copies each) and the whole DiT is trained end to end.
Training checkpoint: step 10,480 (end of training).
Optimizer state: removed (weights only).
Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/mlp-moe/)
Package contents
The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).
| chunk | size | sha256 |
|---|---|---|
mlp-moe.tar.gz.part-00 |
21.47 GB | ab1e1388feb391f19ac711030c4175597feb2e79f58a82508fc96cb80e7155f2 |
mlp-moe.tar.gz.part-01 |
8.30 GB | f613d6eeb910071d8998d3bcccb7e1c908ab19f752c7c2a35ce384c00e11b797 |
Total: 2 chunks, 29.77 GB compressed / 37.14 GB extracted.
Reassemble and verify:
huggingface-cli download Massyzs/mlp-moe --local-dir mlp-moe-package # after your access request is approved
cd mlp-moe-package
cat mlp-moe.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/ # creates <BASE_DIR>/ckpt/mlp-moe/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents) # optional check
MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).
Extracted files:
| file | size | sha256 |
|---|---|---|
mlp-moe/meta.json |
373 B | d0a7b3e850616b946d2ebb23b50514136502c74feb30e1c75b99265a74b07542 |
mlp-moe/moe_v3_trainable.safetensors |
37.14 GB | 82bfbe6509636b5027385ecb92c1168623bdd4fe33532b02bfe7dbe5439d8fe3 |
Base weights you also need
The package contains the full DiT (18.57B parameters, bf16); only the MLLM encoder (qwen/), the VAE and moe_expert_init.safetensors are needed in addition.
All frozen base components are published, already converted to the layout this code expects, in
Massyzs/kiwi-edit-5b-instruct-only-videoxfun:
huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun
See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.
Inference
python scripts/moe/mlp-moe/moe_v3_infer.py --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/mlp-moe \
--src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"
Single 80 GB GPU. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, segmented text CFG (scale 3 on the first 60% of the steps). Output: <OUT_DIR>/<name>.mp4, lossless PNG frames in <OUT_DIR>/frames/<name>/, and per-layer expert-routing maps.
Architecture
30-block DiT (dim 3072); in blocks 1,3,...,27 the FFN is a 12-expert MoE (each expert = the original 3072->14336->3072 FFN), top-3 with renormalised gate weights, balance loss alpha 1e-4. Full state dict: 1,457 tensors, 37.1 GB bf16.
Training data
OpenVE single-instruction edits (stage A) followed by multi-instruction data (goku composite + CoinVE, stage B); 10,480 steps, effective batch 192 (8 x 24 GPUs).
License
The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.
- Downloads last month
- -
Model tree for Massyzs/mlp-moe
Base model
linyq/kiwi-edit-5b-instruct-only-diffusers