3D VAE Step2

图像与视频混合训练的左目生成右目,训练至 14,000 steps 的原始 full checkpoint。

代码与完整准备流程:WangHaiyan378/3D_VAE,发布时对应提交 77f53f5。

  • 权重文件:step_14000.pt,17,030,129,902 bytes。
  • Wan2.1-T2V-1.3B DiT,全量微调;32 通道输入由噪声 latent(16)与左目 latent(16)拼接。
  • checkpoint 包含 model、optim、step、finetune、cfg,保留完整优化器状态。
  • 配套 stage2_mix.yaml 是代码仓库的可迁移配置模板;原始训练配置保存在 checkpoint 的 cfg 中,包含原训练环境路径,使用时请按本机路径调整。
  • 本仓库只提供该阶段 checkpoint;官方 Wan/VAE/T5 权重和文本缓存需要按代码仓库说明准备。
  • 文件是项目自定义 PyTorch checkpoint,需要先安装并在代码仓库内使用;不是 Diffusers 或 Transformers 的 from_pretrained 格式。原始配置包含项目 Config 对象,项目加载器使用 weights_only=False。

下载及使用

在代码仓库根目录、安装依赖后执行:

from huggingface_hub import hf_hub_download
from pathlib import Path
import shutil

downloaded = hf_hub_download('sdfwasdf/3D_VAE_Step2_model', 'step_14000.pt')
target = Path('output/mix_1.3b_full/checkpoints/latest.pt')
target.parent.mkdir(parents=True, exist_ok=True)
shutil.copyfile(downloaded, target)

按代码仓库 README 准备官方权重与文本缓存,再运行:

python -m video.stereo_mix.infer.infer_video --config configs/stage2_mix.yaml --ckpt output/mix_1.3b_full/checkpoints/latest.pt --input left.mp4 --output right.mp4 --num-frames 81 --steps 30

本阶段从 Step1 的 step_4000.pt 初始化。可直接作为代码仓库 Stage3 深度条件训练的 model.init_from。

文件字节数与 SHA-256 见 checkpoint_info.json 和 SHA256SUMS。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for sdfwasdf/3D_VAE_Step2_model

Finetuned
(99)
this model