You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Access to these weights is granted manually by the authors. Please tell us your name, affiliation and intended use.
Log in or Sign Up to review the conditions and access this model content.
videoeditor-14b (Kiwi-Edit 5B with 7-expert FFN + attention MoE)
The Kiwi-Edit 5B DiT where, in 14 of its 30 blocks, both the FFN and the self-attention key/value projections are mixtures of 7 full-weight experts (the original weights, five category-specialised experts and one direct fine-tune), top-3 routing with a log-gate prior injected into attention, no shared expert. 13.99B parameters in total, ~8B active per token.
Training checkpoint: step 18,630 (end of training, one epoch).
Optimizer state: removed (weights only).
Code: https://github.com/massyzs/Video-Editor-MoE (scripts/moe/videoeditor-14b/)
Package contents
The weights are shipped as a tar.gz stream split into ~20 GB chunks (Hugging Face limits single files to 50 GB).
| chunk | size | sha256 |
|---|---|---|
videoeditor-14b.tar.gz.part-00 |
21.47 GB | 7e51c27d3102d5da475dd9f906d5e6902fada5a43960ba3decfd4bfb7db4ab6c |
videoeditor-14b.tar.gz.part-01 |
0.95 GB | 11e34da0e0c4657b68313a5e8fd8c1dcb8695f40c6569d3184b515d3eb82bceb |
Total: 2 chunks, 22.42 GB compressed / 27.97 GB extracted.
Reassemble and verify:
huggingface-cli download Massyzs/videoeditor-14b --local-dir videoeditor-14b-package # after your access request is approved
cd videoeditor-14b-package
cat videoeditor-14b.tar.gz.part-* | tar -xzf - -C <BASE_DIR>/ckpt/ # creates <BASE_DIR>/ckpt/videoeditor-14b/
cp SHA256SUMS.contents <BASE_DIR>/ckpt/ && (cd <BASE_DIR>/ckpt && sha256sum -c SHA256SUMS.contents) # optional check
MANIFEST.json lists every payload file with its size and sha256, and every chunk with its sha256 (SHA256SUMS.parts).
Extracted files:
| file | size | sha256 |
|---|---|---|
videoeditor-14b/meta.json |
455 B | 76c0486df3e6824f58ff3033a669d6fec1efcf148e9b94b6bfb369b2f80d9c04 |
videoeditor-14b/moe_v3_1_trainable.safetensors |
27.97 GB | b7dea5b5ccf708eea36376243099f175bc03272748fe854b4c49b79e14ae8e5a |
Base weights you also need
The package contains the full DiT (13.99B parameters, bf16); only the MLLM encoder (qwen/), the VAE and moe_expert_init.safetensors are needed in addition.
All frozen base components are published, already converted to the layout this code expects, in
Massyzs/kiwi-edit-5b-instruct-only-videoxfun:
huggingface-cli download Massyzs/kiwi-edit-5b-instruct-only-videoxfun --local-dir <BASE_DIR>/ckpt/kiwi-edit-instruct-only-videoxfun
See BASE_WEIGHTS.md for where each component comes from (Kiwi-Edit, Wan2.2) and its checksum.
Inference
python scripts/moe/videoeditor-14b/moe_v3_1_infer.py --base_dir <BASE_DIR> --ckpt <BASE_DIR>/ckpt/videoeditor-14b \
--src_video input.mp4 --prompt "Replace the red car with a blue truck" --name edited --out_dir <OUT_DIR>
# multi-step instruction: --prompts "Remove the dog ||| Convert to oil painting style"
Single 80 GB GPU. Defaults: 50 steps, 49 frames at 480x832, sigma shift 5, segmented text CFG (scale 3 on the first 60% of the steps). Output: <OUT_DIR>/<name>.mp4, lossless PNG frames in <OUT_DIR>/frames/<name>/, and per-layer expert-routing maps.
Architecture
In blocks 1,3,...,27: FFN MoE (7 experts, top-3) and self-attention k/v MoE (7 (k,v) expert pairs sharing one softmax gate, top-3, selected keys/values concatenated with a log-gate prior). Full state dict: 1,527 tensors, 28.0 GB bf16.
Training data
OpenVE single-instruction edits mixed with multi-instruction data (goku composite + CoinVE, x8 repeat); 18,630 steps = 1 epoch, effective batch 192.
License
The released weights and code are under Apache-2.0. The frozen base components (Kiwi-Edit MLLM encoder / DiT base, Wan2.2 VAE) are not redistributed here and remain subject to their own licenses.
- Downloads last month
- -
Model tree for Massyzs/videoeditor-14b
Base model
linyq/kiwi-edit-5b-instruct-only-diffusers